← Blog/blog/eos-token-semantic-stop

Two stop tokens can teach a model never to stop

01

The decoder sees two exits; the loss sees two labels

An end-of-sequence token, or EOS token, tells a language model to stop generating. Modern model families can have several: one may end raw text, another an assistant turn, and another a message. In a single-turn math rollout, two different surface tokens can perform the same operational action: stop.

Yang and coauthors show that sampled-token on-policy distillation can miss that equivalence. The student generates from its own policy; the fixed teacher scores only the token actually sampled. If the student ends with token A while the teacher puts its stopping probability on token B, the update can punish A even though both models want to stop.

02

One log ratio suppresses a perfectly good stop

At prefix s, sampled-token distillation assigns the sampled token y the detached coefficient A = log pteacher(y|s) − log pstudent(y|s). A negative coefficient suppresses that sampled action. The paper's Qwen3 base student initially puts roughly 0.8 probability on its native EOS, while its probability on the teacher's preferred alternative is around 10−11 near terminal states.

In the illustrative split used here, the student has EOS masses 0.8 and 10−11; the teacher has 0.01 and 0.79. Their total stopping masses are 0.80 and 0.80. Yet the native-token coefficient is -4.382, while the semantic coefficient is approximately 0.000. Same decision, opposite training story.

import math

def stopping_mass(probabilities):
    total = sum(probabilities)
    if not probabilities or total < 0 or total > 1:
        raise ValueError('invalid probability mass')
    return total

def token_advantage(teacher_probability, student_probability):
    if teacher_probability <= 0 or student_probability <= 0:
        raise ValueError('probabilities must be positive')
    return math.log(teacher_probability) - math.log(student_probability)

def semantic_stop_advantage(teacher_eos, student_eos):
    return token_advantage(stopping_mass(teacher_eos), stopping_mass(student_eos))
native-token advantagesemantic-stop advantageneutral update
Fig 1. Illustrative deterministic probability sweep. The teacher keeps 0.79 mass on its alternative EOS while its mass on the student's native EOS varies. Token-level supervision stays suppressive; semantic aggregation evaluates the total decision to stop. These are formula outputs, not paper measurements.
03

A supported token and a registered token are different

Registering token B as a legal stop helps only after the student samples it. At probability 10−11, the expected wait is 100,000,000,000 draws. The main experiment uses 64 trajectories per step for 200 steps: 12,800 trajectories before accounting for their many nonterminal token positions. Even 12,800 independent chances at exactly that probability give only about 1.28 × 10−7 chance of seeing the token once.

Fig 2. Chance of sampling a 10⁻¹¹-probability token at least once, scaled by one billion for legibility. The four budgets are illustrative independent-draw calculations; 64 and 12,800 mirror the paper's trajectories per step and across 200 main steps, not the exact number of eligible terminal prefixes.
04

Collapse a token class, not the vocabulary

The cleanest cross-family correction defines one semantic action, STOP, whose probability is the sum over a declared equivalence class of EOS tokens. If the sampled token is in that class, both teacher and student are scored on total stopping mass. Non-EOS tokens are unchanged. No canonical spelling of “stop” must win.

InterventionWhat changesPaper result
Shared-set decodingdecoder onlytracks vanilla failure
Teacher-side mappingteacher EOS massmitigates inflation
Semantic EOS classobjective's stop actionmitigates across 3 families
Canonical single EOSteacher + student action spacemitigates inflation
The paper's Qwen3 ablation separates decoder configuration from probability-level supervision; semantic aggregation is then tested across model families.
FamilyBase preferencePost-trained preferenceDeclared EOS sets
Qwen3<|endoftext|><|im_end|> (plus <|endoftext|>)different sets
Llama 3.2<|end_of_text|><|eom_id|> / <|eot_id|>different sets
Gemma 3<eos><end_of_turn>same set
Condensed from the paper's token table and probability analysis. Gemma is the key control: the declared sets already match, but the learned surface preferences do not.
05

A small stopping error becomes a giant length bill

To build intuition, suppose each token position has a constant stop probability q. Then generation length is geometric, with untruncated mean 1/q. The real model's q changes with its prefix, so this is not a reproduction of training. It is a transparent bridge from falling stopping mass to longer responses and budget clipping.

expected tokensclipping probability × 100
Fig 3. Seed-free geometric toy model with a 100-token budget. Expected truncated length and clipping probability are computed by the core module. Moving left to right means q falls from 0.5 to 0.01; the mixed axis is intentional because both quantities are on a 0–100 scale.

The paper's actual budget is much larger: up to 7,168 response tokens during training. Under vanilla Qwen3 distillation, the native EOS probability falls from roughly 0.8 to near zero while response length and clipping rise. The same qualitative termination collapse appears in all three studied families.

06

The correction works—and then the failure returns

That residual is the paper's most important boundary. Semantic aggregation fixes a specific support-and-supervision bug and is largely inert when student and teacher termination representations already align. It does not remove every continuation bias, entropy collapse, or degradation of teacher guidance on long, repetitive prefixes.

The evidence is also single-turn mathematical reasoning. In an agentic system, “end assistant turn,” “end document,” and “handoff to tool” may change control flow and are not automatically equivalent. The EOS class must be defined by protocol semantics, not by token names alone.

07

Evaluation has its own termination token

The paper catches a second interface trap in its appendix. The original DAPO grader scores the Qwen3-4B teacher at 6.28% Avg@16; a parser that accepts equivalent multiline LaTeX layouts raises it to 24.34%. Several students can appear to beat the teacher under the brittle parser. A model can terminate correctly and still be declared wrong by an answer extractor that expects a different surface form.

The operational checklist is short: audit probability mass across every valid stop token at the same prefixes; align the objective, not only the decoder; test the correction where no mismatch exists; and keep monitoring late-stage length after the early bug disappears. For the architecture underneath those token probabilities, continue with the Transformer walkthrough or the neural-network walkthrough.

References

  1. Y. Yang, T. Yu, S. Li, K. Zhao, X. Zhang, C. Bansal, H. Yao, T. W. Killian, and W. Zhang (2026). When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation. arXiv:2609.20511