← Blog/blog/conformal-recommendation-selection-gap

The calibrated recommender can still pick the wrong item

A recommendation system can keep a good item in the candidate set and still make a bad recommendation. Si and Wei's RLCP method tackles the first half: it learns a variable-size action set and adjusts a threshold online so a proxy target is missed at a chosen rate. The paper's own decomposition makes the second half impossible to ignore.

01

Replace a fixed slate with a calibrated threshold

Sequential recommendation optimizes a session, not just one click. Most systems still pass a fixed number of items downstream. RLCP instead scores every candidate by its estimated gap from the critic's best action and retains items whose gap is below a threshold. A cap limits the resulting set, and the smallest-gap item is a fallback when nothing passes.

The threshold reacts to a proxy miss: whether the raw set contains at least one action labeled near-optimal by an observable proxy. A miss raises the threshold; a hit lowers it. That feedback loop is the conformal part of the method.

def calibrate(margins, alpha, eta, threshold):
    history = []
    misses = 0
    for step, margin in enumerate(margins, 1):
        miss = int(margin > threshold)
        misses += miss
        next_threshold = threshold + eta * (miss - alpha)
        history.append((step, threshold, miss, misses / step))
        threshold = next_threshold
    return history

The site runs the TypeScript translation above. The following deterministic arithmetic toy repeats ten fixed proxy margins for 80 steps; it is not a reproduction of the paper's learned recommender. The threshold expands after a miss and contracts after a hit.

proxy marginonline threshold
Illustrative deterministic margins and the online threshold for α=0.20 and η=0.18. Every point is computed by the core function shown above.
running proxy miss rate20% target
The realized proxy miss rate is steered toward 20%; after 80 illustrative interactions it is 22.5%. This is pathwise accounting, not a claim about unseen users.
02

A threshold still becomes a capped set

Raising the threshold makes the raw set less selective. Once a maximum size is imposed, however, a proxy-worthy item can pass the threshold and then be removed by truncation. Calibration controls raw misses; ranking and capacity control these extra capped misses.

executed set size (cap 4)
Illustrative executed-set size as the threshold sweeps across six critic gap scores. The fallback keeps size at one; the cap stops growth at four.
03

Filtering loss and selection loss are different debts

Let the best unrestricted action value be one. Filtering loss is the gap between that value and the best value left in the set. Selection loss is the remaining gap between the best retained value and what the downstream selector actually chooses. The paper proves that discounted value loss is the occupancy-weighted sum of both, scaled by 1/(1−γ).

One-state arithmetic with optimal values [1.00, 0.65, 0.20]. Perfect retention can still lose 0.80 when the selector chooses the worst retained item.
04

The observable proxy may rank the wrong future

RLCP cannot observe the unrestricted optimal action value. In the experiments, the proxy target comes from the simulator's immediate click or like probability. That is useful feedback, but an action can produce a slightly smaller immediate response and a much better next state. The paper explicitly carries this continuation error into its conditional value bound.

item A: proxy winneritem B: continuation winner
Illustrative two-item reversal. Item A has immediate reward 0.90 and no continuation; B has 0.80 plus continuation value 1. At γ=1, following the immediate proxy leaves 0.90 long-term value on the table.
LayerWhat it controlsWhat still needs evidence
Online updateObserved proxy misses along the realized trajectoryFuture-state or population coverage
Proxy retentionA proxy-worthy item remains in the raw setWhether it has high long-term optimal value
Action-set hitAt least one target item survives filteringWhether the selector chooses that item
Set-size capMaximum alternatives passed downstreamA target item surviving truncation
Simulator resultPerformance under its response and leave modulesReal-user counterfactual feedback
The paper states these boundaries explicitly; the middle row is the fragile bridge from calibration to long-term value.
05

What the experiments actually establish

The paper evaluates KuaiRand-Pure and MovieLens 1M through the KuaiSim simulator, comparing two RLCP variants with A2C, DDPG, TD3, and HAC. The headline is exposure: in every one of 19 configurations, at least one RLCP variant has the highest catalog diversity while respecting the same maximum set size. Session depth is competitive, not uniformly best.

Paper-reported scopeConfigsCatalog / within-list diversitySet-size result
All experiments19Best RLCP variant leads in every configurationNo larger than matched baseline cap
KuaiRand-Pure10RLCP-Single leads 10/10; 1.48×–4.56×Smaller in 9/10; up to 25.5% smaller
MovieLens 1M9Better variant gives 1.11×–5.21×Full RLCP smaller in 4 configurations
Full RLCP19Within-list diversity reported as 1.00Capped at the baseline slate size
Paper-reported summaries. Bold claims compare point estimates; the appendix says they do not establish statistical significance.

The simulator also provides the proxy-target verifier: it knows the Bernoulli success probability for every candidate. A real platform usually observes feedback only for exposed items. The paper notes that executed-item feedback alone cannot identify whether an unchosen target was present, so deployment needs a counterfactual response model or another verifier.

06

What to probe next

First, report raw proxy misses, cap-induced misses, and final chosen-item regret separately. Second, stress-test the method when the response proxy and long-horizon value disagree by construction. Third, replace the simulator's all-item verifier with logged bandit feedback and quantify the identification error. Finally, attach uncertainty to the 19 configuration comparisons rather than ranking point estimates alone.

References

  1. Wenwen Si and Honghao Wei (2026). Reinforcement Learning with Conformal Action Sets: An Application to Sequential Recommendation. arXiv:2610.08743
  2. Isaac Gibbs and Emmanuel Candès (2021). Adaptive Conformal Inference Under Distribution Shift. NeurIPS 2021
  3. Kaili Zhao et al. (2023). KuaiSim: A Comprehensive Simulator for Recommender Systems. NeurIPS 2023 Datasets and Benchmarks Track