Sparse Mixture-of-Experts models buy capacity by routing each token through only a few feed-forward networks, called experts. Jha, Li, Leskovec, Liang, and Zettlemoyer ask what happens when scarce text is replayed. Across models with 80 million to one billion active parameters, the sparse models overfit earlier than dense Transformers and lose their all-unique-data advantage at high repetition.
The memorable headline is a 32-pass reversal. The more important reading is underneath it: the experiment fixes total training tokens, so increasing the repetition rate necessarily removes unique tokens. “Four repeats” is therefore not an independent treatment or a universal epoch limit; it is one point on a unique-data-per-parameter curve.
Two kinds of sparsity meet one shrinking corpus
An MoE router scores every expert but activates only the top k. The layer owns many total parameters, yet each token touches far fewer active parameters. Data repetition creates the mirror image: training consumes many total tokens, but some are copies of the same unique tokens.
The study holds compute exposure at T = 20Na, where Na is active parameter count. With repetition R, the unique pool is U = T/R. This small identity is the paper's experimental axis and our reconstruction:
def repetition_budget(active, total, tokens_per_active, repeats):
train_tokens = active * tokens_per_active
unique_tokens = train_tokens / repeats
return {
'train_tokens': train_tokens,
'unique_tokens': unique_tokens,
'unique_per_active': unique_tokens / active,
'unique_per_total': unique_tokens / total,
}
def expert_unique_exposure(unique_tokens, experts, active_experts):
return unique_tokens * active_experts / expertsThe page runs the TypeScript implementation above; Python and C++ are line-for-line translations. No model training is simulated in these charts. They expose the exact budget arithmetic behind the paper's real training runs.
Why one expert sees an even smaller world
The paper's mechanism starts with router ossification: token-to-expert assignments become stable early. Under a balanced, fixed top-4-of-64 router, each expert receives only 4/64 of the unique-token assignments on average. At R = 32, the 50M-token unique pool becomes just 3.125M unique token-exposures per expert under this accounting approximation.
This is not merely a load-balancing story. The authors find that load imbalance and the auxiliary balancing loss do not track repetition consistently. What does rise is expert specialization, measured by how much held-out cross-entropy increases when one expert is knocked out. From R = 1 to 32, the median knockout effect grows 1.1× with 16 experts and 2.3× with 128 experts.
Routing freezes early, but that is correlation—not verdict
Routing stability is the fraction of validation tokens that keep the same top-1 expert between consecutive checkpoints, computed per layer and then averaged. The definition is simple enough to audit:
def top1_stability(previous, current):
layer_scores = []
for before, after in zip(previous, current):
unchanged = sum(a == b for a, b in zip(before, after))
layer_scores.append(unchanged / len(before))
return sum(layer_scores) / len(layer_scores)Stable routing supports the “same small shard, again and again” story, but it is not a causal certificate. The paper itself reports that router jitter does not improve performance, while dropout and output masking do. Dropout sharply reduces overfitting and expert knockout cost without restoring substantial router plasticity. The likely failure sits in what fixed experts learn, not simply in the fact that routes become fixed.
| Comparison | Paper-reported result | Scope |
|---|---|---|
| 80M dense | minimal degradation through 8× | held-out CE |
| 80M-active MoE | noticeable degradation at 4× | held-out CE |
| Dense versus MoE | ranking reverses at 32× | matched active parameters |
| Strong masking | MoE still beats dense beyond 64× | regularized runs |
| No intervention | recovers all-unique performance | none tested |
The regularizer points away from the router
| Intervention | Repetition effect | What it perturbs |
|---|---|---|
| Residual dropout | large mitigation | sub-layer outputs |
| FFN / expert masking | large mitigation | parameter outputs |
| Weight decay | within noise | weight magnitude |
| Gradient clipping | within variance | update norm |
| Router jitter | no clear impact | routing inputs |
Strong residual dropout, FFN output masking, expert dropout, and expert output masking all help. Weight decay, clipping, and router jitter do not show measurable gains in these runs. That pattern is consistent with breaking co-adapted expert functions. It does not prove a unique mechanism: dropout changes optimization and representation learning in many ways.
What to probe next
A decisive follow-up would vary repetition while holding the unique pool fixed by changing total optimization steps, then separately vary U at fixed R. Cross those axes with total parameters, router plasticity, and expert dropout. Report per-expert unique coverage—not only average load—and measure whether the same token clusters stay assigned after domain shifts.
The present evidence spans several scales and domains, but still uses one training recipe, nested data subsets from a shared permutation, fixed active-parameter matching, and cross-entropy-centered evaluation. The authors show seed sensitivity and downstream tasks, yet the numeric breakpoints remain recipe-specific. For the routing computation itself, continue with the site's Transformer walkthrough; for the expert networks and dropout, see the neural-network page.
References
- Jha, Atindra; Li, Margaret; Leskovec, Jure; Liang, Percy; and Zettlemoyer, Luke (2026). Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data. arXiv:2609.11917
- Shazeer, Noam et al. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. ICLR 2017
- Muennighoff, Niklas et al. (2023). Scaling Data-Constrained Language Models. NeurIPS 2023