The decoder sees two exits; the loss sees two labels
An end-of-sequence token, or EOS token, tells a language model to stop generating. Modern model families can have several: one may end raw text, another an assistant turn, and another a message. In a single-turn math rollout, two different surface tokens can perform the same operational action: stop.
Yang and coauthors show that sampled-token on-policy distillation can miss that equivalence. The student generates from its own policy; the fixed teacher scores only the token actually sampled. If the student ends with token A while the teacher puts its stopping probability on token B, the update can punish A even though both models want to stop.
One log ratio suppresses a perfectly good stop
At prefix s, sampled-token distillation assigns the sampled token y the detached coefficient A = log pteacher(y|s) − log pstudent(y|s). A negative coefficient suppresses that sampled action. The paper's Qwen3 base student initially puts roughly 0.8 probability on its native EOS, while its probability on the teacher's preferred alternative is around 10−11 near terminal states.
In the illustrative split used here, the student has EOS masses 0.8 and 10−11; the teacher has 0.01 and 0.79. Their total stopping masses are 0.80 and 0.80. Yet the native-token coefficient is -4.382, while the semantic coefficient is approximately 0.000. Same decision, opposite training story.
import math
def stopping_mass(probabilities):
total = sum(probabilities)
if not probabilities or total < 0 or total > 1:
raise ValueError('invalid probability mass')
return total
def token_advantage(teacher_probability, student_probability):
if teacher_probability <= 0 or student_probability <= 0:
raise ValueError('probabilities must be positive')
return math.log(teacher_probability) - math.log(student_probability)
def semantic_stop_advantage(teacher_eos, student_eos):
return token_advantage(stopping_mass(teacher_eos), stopping_mass(student_eos))A supported token and a registered token are different
Registering token B as a legal stop helps only after the student samples it. At probability 10−11, the expected wait is 100,000,000,000 draws. The main experiment uses 64 trajectories per step for 200 steps: 12,800 trajectories before accounting for their many nonterminal token positions. Even 12,800 independent chances at exactly that probability give only about 1.28 × 10−7 chance of seeing the token once.
Collapse a token class, not the vocabulary
The cleanest cross-family correction defines one semantic action, STOP, whose probability is the sum over a declared equivalence class of EOS tokens. If the sampled token is in that class, both teacher and student are scored on total stopping mass. Non-EOS tokens are unchanged. No canonical spelling of “stop” must win.
| Intervention | What changes | Paper result |
|---|---|---|
| Shared-set decoding | decoder only | tracks vanilla failure |
| Teacher-side mapping | teacher EOS mass | mitigates inflation |
| Semantic EOS class | objective's stop action | mitigates across 3 families |
| Canonical single EOS | teacher + student action space | mitigates inflation |
| Family | Base preference | Post-trained preference | Declared EOS sets |
|---|---|---|---|
| Qwen3 | <|endoftext|> | <|im_end|> (plus <|endoftext|>) | different sets |
| Llama 3.2 | <|end_of_text|> | <|eom_id|> / <|eot_id|> | different sets |
| Gemma 3 | <eos> | <end_of_turn> | same set |
A small stopping error becomes a giant length bill
To build intuition, suppose each token position has a constant stop probability q. Then generation length is geometric, with untruncated mean 1/q. The real model's q changes with its prefix, so this is not a reproduction of training. It is a transparent bridge from falling stopping mass to longer responses and budget clipping.
The paper's actual budget is much larger: up to 7,168 response tokens during training. Under vanilla Qwen3 distillation, the native EOS probability falls from roughly 0.8 to near zero while response length and clipping rise. The same qualitative termination collapse appears in all three studied families.
The correction works—and then the failure returns
That residual is the paper's most important boundary. Semantic aggregation fixes a specific support-and-supervision bug and is largely inert when student and teacher termination representations already align. It does not remove every continuation bias, entropy collapse, or degradation of teacher guidance on long, repetitive prefixes.
The evidence is also single-turn mathematical reasoning. In an agentic system, “end assistant turn,” “end document,” and “handoff to tool” may change control flow and are not automatically equivalent. The EOS class must be defined by protocol semantics, not by token names alone.
Evaluation has its own termination token
The paper catches a second interface trap in its appendix. The original DAPO grader scores the Qwen3-4B teacher at 6.28% Avg@16; a parser that accepts equivalent multiline LaTeX layouts raises it to 24.34%. Several students can appear to beat the teacher under the brittle parser. A model can terminate correctly and still be declared wrong by an answer extractor that expects a different surface form.
The operational checklist is short: audit probability mass across every valid stop token at the same prefixes; align the objective, not only the decoder; test the correction where no mismatch exists; and keep monitoring late-stage length after the early bug disappears. For the architecture underneath those token probabilities, continue with the Transformer walkthrough or the neural-network walkthrough.
References
- Y. Yang, T. Yu, S. Li, K. Zhao, X. Zhang, C. Bansal, H. Yao, T. W. Killian, and W. Zhang (2026). When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation. arXiv:2609.20511