A new policy can already be hiding between old ones
Reinforcement learning from a reward normally means another expensive training run. Hamidieh, Daras, and Torralba ask a sharper question: if several policies have already been optimized from the same reference model, can their outputs predict the policy a new reward would produce? Their method, PoEM, composes the old policies at inference time—no parameter update and no samples from the target policy are required.
The idea follows from KL-regularized RL. KL divergence penalizes moving too far from a reference policy. With reward r and penalty β, the exact optimum is the reference probability multiplied by exp(r / β), then normalized. A weighted sum of rewards therefore becomes a weighted product of policy-to-reference ratios.
import math
def compose(reference, experts, weights, gamma=1.0):
logits = []
for outcome, p_ref in enumerate(reference):
ref_log = math.log(p_ref)
shift = sum(
weight * (math.log(expert[outcome]) - ref_log)
for expert, weight in zip(experts, weights)
)
logits.append(ref_log + gamma * shift)
maximum = max(logits)
scores = [math.exp(value - maximum) for value in logits]
total = sum(scores)
return [score / total for score in scores]Coverage is the real gate
A policy's log ratio log(πₖ / πref) is its implicit reward direction, up to a prompt-only constant. PoEM regresses the new reward scores on these directions. The paper calls held-out R² from that regression coverage: how much of the requested reward variation the existing policies can express.
Our toy has two orthogonal directions. A target halfway between them has coverage 1.00. The vector [1, 1, −1, −1] is perpendicular to both and has coverage 0.00. Scaling the fitted mixture with γ changes its length, never its direction.
Held-out experts expose the missing directions
The cleanest test removes one expert from a basis, treats it as an unknown target, and fits only from its reward scores. Across programmatic GRPO, programmatic DPO, and diverse learned reward models, coverage correlates with reward error at Spearman ρ from −0.71 to −0.81. The direction diagnostic predicts when composition will fail.
Matching reward is not matching the policy
| Basis | Targets | PoEM reward error | Top expert | Best-of-16 | Relative KL | PoEM closer |
|---|---|---|---|---|---|---|
| P-GRPO | 32 | 0.28 | 0.23 | 0.37 | 0.24 | 32 / 32 |
| P-DPO | 32 | 0.19 | 0.32 | 0.40 | 1.05 | 13 / 32 |
| RM-RB | 8 | 0.05 | 0.13 | 0.24 | 0.29 | 8 / 8 |
| RM-HH | 3 | 0.07 | 0.07 | 0.22 | 0.58 | 3 / 3 |
P-DPO is the revealing row. PoEM improves reward error over the top expert, 0.19 versus 0.32, yet its relative KL is 1.05 and it is closer to the trained target policy for only 13 of 32 rewards. The DPO experts were trained offline, and the ideal exponential-tilt derivation need not describe their actual update. A reward score can agree while the output distribution does not.
Autoregressive composition adds another approximation
For a transformer, sequence probabilities factor token by token. PoEM adds each expert's next-token log-probability minus the reference log-probability, then applies a local softmax. This is practical, but local normalization is not generally identical to one global normalization over complete sequences: continuation normalizers can depend on the prefix.
The derivation also assumes a shared reference policy and regularization strength. Finite models and imperfect optimization produce implicit reward directions that may differ from the named rewards. This is why the empirical coverage regression uses observed expert log ratios, not just a reward catalog. The related DPO walkthrough shows why changing the reference changes the preference objective itself.
What to probe next
Before composing, estimate coverage on held-out prompts and report it beside reward recovery, KL to any available target expert, and output diversity. Add basis policies for low-coverage target clusters rather than turning γ upward. Then test whether the composed policy is a useful warm start for a short target RL run or can be distilled into one model; the paper identifies both directions but does not evaluate them.
Cost matters too. PoEM avoids retraining but runs every basis expert at inference. The experiments use small models—Qwen3-0.6B adapters for text and Stable Diffusion 1.4 adapters for images—so quality and latency at larger scale remain open. The honest promise is narrower than zero-shot RL: reuse a well-covered policy basis, diagnose when it is missing a direction, and decline to extrapolate when it is.
References
- K. Hamidieh, G. Daras, and A. Torralba (2026). PoEM: Predicting RL Outcomes from Existing Policies. arXiv:2609.30226 [cs.LG]
- R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. Advances in Neural Information Processing Systems 36
- L. Schulman, J. Chen, and P. Abbeel (2017). Equivalence Between Policy Gradients and Soft Q-Learning. arXiv:1704.06440