A recommendation system can keep a good item in the candidate set and still make a bad recommendation. Si and Wei's RLCP method tackles the first half: it learns a variable-size action set and adjusts a threshold online so a proxy target is missed at a chosen rate. The paper's own decomposition makes the second half impossible to ignore.
Replace a fixed slate with a calibrated threshold
Sequential recommendation optimizes a session, not just one click. Most systems still pass a fixed number of items downstream. RLCP instead scores every candidate by its estimated gap from the critic's best action and retains items whose gap is below a threshold. A cap limits the resulting set, and the smallest-gap item is a fallback when nothing passes.
The threshold reacts to a proxy miss: whether the raw set contains at least one action labeled near-optimal by an observable proxy. A miss raises the threshold; a hit lowers it. That feedback loop is the conformal part of the method.
def calibrate(margins, alpha, eta, threshold):
history = []
misses = 0
for step, margin in enumerate(margins, 1):
miss = int(margin > threshold)
misses += miss
next_threshold = threshold + eta * (miss - alpha)
history.append((step, threshold, miss, misses / step))
threshold = next_threshold
return historyThe site runs the TypeScript translation above. The following deterministic arithmetic toy repeats ten fixed proxy margins for 80 steps; it is not a reproduction of the paper's learned recommender. The threshold expands after a miss and contracts after a hit.
A threshold still becomes a capped set
Raising the threshold makes the raw set less selective. Once a maximum size is imposed, however, a proxy-worthy item can pass the threshold and then be removed by truncation. Calibration controls raw misses; ranking and capacity control these extra capped misses.
Filtering loss and selection loss are different debts
Let the best unrestricted action value be one. Filtering loss is the gap between that value and the best value left in the set. Selection loss is the remaining gap between the best retained value and what the downstream selector actually chooses. The paper proves that discounted value loss is the occupancy-weighted sum of both, scaled by 1/(1−γ).
The observable proxy may rank the wrong future
RLCP cannot observe the unrestricted optimal action value. In the experiments, the proxy target comes from the simulator's immediate click or like probability. That is useful feedback, but an action can produce a slightly smaller immediate response and a much better next state. The paper explicitly carries this continuation error into its conditional value bound.
| Layer | What it controls | What still needs evidence |
|---|---|---|
| Online update | Observed proxy misses along the realized trajectory | Future-state or population coverage |
| Proxy retention | A proxy-worthy item remains in the raw set | Whether it has high long-term optimal value |
| Action-set hit | At least one target item survives filtering | Whether the selector chooses that item |
| Set-size cap | Maximum alternatives passed downstream | A target item surviving truncation |
| Simulator result | Performance under its response and leave modules | Real-user counterfactual feedback |
What the experiments actually establish
The paper evaluates KuaiRand-Pure and MovieLens 1M through the KuaiSim simulator, comparing two RLCP variants with A2C, DDPG, TD3, and HAC. The headline is exposure: in every one of 19 configurations, at least one RLCP variant has the highest catalog diversity while respecting the same maximum set size. Session depth is competitive, not uniformly best.
| Paper-reported scope | Configs | Catalog / within-list diversity | Set-size result |
|---|---|---|---|
| All experiments | 19 | Best RLCP variant leads in every configuration | No larger than matched baseline cap |
| KuaiRand-Pure | 10 | RLCP-Single leads 10/10; 1.48×–4.56× | Smaller in 9/10; up to 25.5% smaller |
| MovieLens 1M | 9 | Better variant gives 1.11×–5.21× | Full RLCP smaller in 4 configurations |
| Full RLCP | 19 | Within-list diversity reported as 1.00 | Capped at the baseline slate size |
The simulator also provides the proxy-target verifier: it knows the Bernoulli success probability for every candidate. A real platform usually observes feedback only for exposed items. The paper notes that executed-item feedback alone cannot identify whether an unchosen target was present, so deployment needs a counterfactual response model or another verifier.
What to probe next
First, report raw proxy misses, cap-induced misses, and final chosen-item regret separately. Second, stress-test the method when the response proxy and long-horizon value disagree by construction. Third, replace the simulator's all-item verifier with logged bandit feedback and quantify the identification error. Finally, attach uncertainty to the 19 configuration comparisons rather than ranking point estimates alone.
References
- Wenwen Si and Honghao Wei (2026). Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation. arXiv:2610.08743
- Isaac Gibbs and Emmanuel Candès (2021). Adaptive Conformal Inference Under Distribution Shift. NeurIPS 2021
- Kaili Zhao et al. (2023). KuaiSim: A Comprehensive Simulator for Recommender Systems. NeurIPS 2023 Datasets and Benchmarks Track