← Blog/blog/policy-composition-coverage-gap

The RL shortcut that cannot invent a missing direction

01

A new policy can already be hiding between old ones

Reinforcement learning from a reward normally means another expensive training run. Hamidieh, Daras, and Torralba ask a sharper question: if several policies have already been optimized from the same reference model, can their outputs predict the policy a new reward would produce? Their method, PoEM, composes the old policies at inference time—no parameter update and no samples from the target policy are required.

The idea follows from KL-regularized RL. KL divergence penalizes moving too far from a reference policy. With reward r and penalty β, the exact optimum is the reference probability multiplied by exp(r / β), then normalized. A weighted sum of rewards therefore becomes a weighted product of policy-to-reference ratios.

import math

def compose(reference, experts, weights, gamma=1.0):
    logits = []
    for outcome, p_ref in enumerate(reference):
        ref_log = math.log(p_ref)
        shift = sum(
            weight * (math.log(expert[outcome]) - ref_log)
            for expert, weight in zip(experts, weights)
        )
        logits.append(ref_log + gamma * shift)
    maximum = max(logits)
    scores = [math.exp(value - maximum) for value in logits]
    total = sum(scores)
    return [score / total for score in scores]
outcome 1outcome 2outcome 3outcome 4
Fig 1. Seedless four-outcome toy computed by the core implementation. Moving α from 0 to 1 composes two exact KL-regularized experts; each line is one outcome probability. The direct target-policy calculation matches the composition to floating-point precision.
02

Coverage is the real gate

A policy's log ratio log(πₖ / πref) is its implicit reward direction, up to a prompt-only constant. PoEM regresses the new reward scores on these directions. The paper calls held-out R² from that regression coverage: how much of the requested reward variation the existing policies can express.

Our toy has two orthogonal directions. A target halfway between them has coverage 1.00. The vector [1, 1, −1, −1] is perpendicular to both and has coverage 0.00. Scaling the fitted mixture with γ changes its length, never its direction.

covered targetmissing direction
Fig 2. Synthetic KL to the desired target as composition strength γ changes. The covered target is recovered at γ = 1. The missing direction stays at the reference for every γ, so its error is flat. These are illustrative values, not paper results.
03

Held-out experts expose the missing directions

The cleanest test removes one expert from a basis, treats it as an unknown target, and fits only from its reward scores. Across programmatic GRPO, programmatic DPO, and diverse learned reward models, coverage correlates with reward error at Spearman ρ from −0.71 to −0.81. The direction diagnostic predicts when composition will fail.

Fig 3. Paper-reported held-out normalized reward recovery. 'Covered' uses R² ≥ 0.3 for P-GRPO and P-DPO; RM-Div uses its median because every R² is below 0.3. The threshold is therefore a within-basis split, not a universal pass mark.
04

Matching reward is not matching the policy

BasisTargetsPoEM reward errorTop expertBest-of-16Relative KLPoEM closer
P-GRPO320.280.230.370.2432 / 32
P-DPO320.190.320.401.0513 / 32
RM-RB80.050.130.240.298 / 8
RM-HH30.070.070.220.583 / 3
Paper-reported combined-reward results. Lower reward error and relative KL are better; 'PoEM closer' counts targets whose policy is nearer to the target expert than the best single basis expert. Best-of-16 has target-reward access at inference.

P-DPO is the revealing row. PoEM improves reward error over the top expert, 0.19 versus 0.32, yet its relative KL is 1.05 and it is closer to the trained target policy for only 13 of 32 rewards. The DPO experts were trained offline, and the ideal exponential-tilt derivation need not describe their actual update. A reward score can agree while the output distribution does not.

geometry correction
Fig 4. Geometry-only correction from the paper, evaluated for two orthogonal toy experts. Equal weights partially cancel direction length, so γgeom rises to √2 at the middle; it rescales existing geometry but adds no missing coordinate.
05

Autoregressive composition adds another approximation

For a transformer, sequence probabilities factor token by token. PoEM adds each expert's next-token log-probability minus the reference log-probability, then applies a local softmax. This is practical, but local normalization is not generally identical to one global normalization over complete sequences: continuation normalizers can depend on the prefix.

The derivation also assumes a shared reference policy and regularization strength. Finite models and imperfect optimization produce implicit reward directions that may differ from the named rewards. This is why the empirical coverage regression uses observed expert log ratios, not just a reward catalog. The related DPO walkthrough shows why changing the reference changes the preference objective itself.

06

What to probe next

Before composing, estimate coverage on held-out prompts and report it beside reward recovery, KL to any available target expert, and output diversity. Add basis policies for low-coverage target clusters rather than turning γ upward. Then test whether the composed policy is a useful warm start for a short target RL run or can be distilled into one model; the paper identifies both directions but does not evaluate them.

Cost matters too. PoEM avoids retraining but runs every basis expert at inference. The experiments use small models—Qwen3-0.6B adapters for text and Stable Diffusion 1.4 adapters for images—so quality and latency at larger scale remain open. The honest promise is narrower than zero-shot RL: reuse a well-covered policy basis, diagnose when it is missing a direction, and decline to extrapolate when it is.

References

  1. K. Hamidieh, G. Daras, and A. Torralba (2026). PoEM: Predicting RL Outcomes from Existing Policies. arXiv:2609.30226 [cs.LG]
  2. R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. Advances in Neural Information Processing Systems 36
  3. L. Schulman, J. Chen, and P. Abbeel (2017). Equivalence Between Policy Gradients and Soft Q-Learning. arXiv:1704.06440