← Blog/blog/moe-repetition-unique-token-tax

The 32-pass MoE that lost to a dense model

Sparse Mixture-of-Experts models buy capacity by routing each token through only a few feed-forward networks, called experts. Jha, Li, Leskovec, Liang, and Zettlemoyer ask what happens when scarce text is replayed. Across models with 80 million to one billion active parameters, the sparse models overfit earlier than dense Transformers and lose their all-unique-data advantage at high repetition.

The memorable headline is a 32-pass reversal. The more important reading is underneath it: the experiment fixes total training tokens, so increasing the repetition rate necessarily removes unique tokens. “Four repeats” is therefore not an independent treatment or a universal epoch limit; it is one point on a unique-data-per-parameter curve.

01

Two kinds of sparsity meet one shrinking corpus

An MoE router scores every expert but activates only the top k. The layer owns many total parameters, yet each token touches far fewer active parameters. Data repetition creates the mirror image: training consumes many total tokens, but some are copies of the same unique tokens.

The study holds compute exposure at T = 20Na, where Na is active parameter count. With repetition R, the unique pool is U = T/R. This small identity is the paper's experimental axis and our reconstruction:

def repetition_budget(active, total, tokens_per_active, repeats):
    train_tokens = active * tokens_per_active
    unique_tokens = train_tokens / repeats
    return {
        'train_tokens': train_tokens,
        'unique_tokens': unique_tokens,
        'unique_per_active': unique_tokens / active,
        'unique_per_total': unique_tokens / total,
    }

def expert_unique_exposure(unique_tokens, experts, active_experts):
    return unique_tokens * active_experts / experts

The page runs the TypeScript implementation above; Python and C++ are line-for-line translations. No model training is simulated in these charts. They expose the exact budget arithmetic behind the paper's real training runs.

80M dense244.4M-total MoE
Fig 1. Exact budget accounting for the paper's 80M-active comparison. Both models consume 1.6B total tokens at every point. The MoE line is lower because its 244.4M total parameters share the same shrinking unique pool. These are computed quantities, not validation losses.
02

Why one expert sees an even smaller world

The paper's mechanism starts with router ossification: token-to-expert assignments become stable early. Under a balanced, fixed top-4-of-64 router, each expert receives only 4/64 of the unique-token assignments on average. At R = 32, the 50M-token unique pool becomes just 3.125M unique token-exposures per expert under this accounting approximation.

dense FFNone balanced MoE expert
Fig 2. A deterministic exposure calculation, not a measured paper curve. A dense FFN sees the full unique pool; a balanced fixed top-4-of-64 expert sees one sixteenth. Real routing is unequal and token overlap makes “unique per expert” a diagnostic, not a new performance claim.

This is not merely a load-balancing story. The authors find that load imbalance and the auxiliary balancing loss do not track repetition consistently. What does rise is expert specialization, measured by how much held-out cross-entropy increases when one expert is knocked out. From R = 1 to 32, the median knockout effect grows 1.1× with 16 experts and 2.3× with 128 experts.

03

Routing freezes early, but that is correlation—not verdict

Routing stability is the fraction of validation tokens that keep the same top-1 expert between consecutive checkpoints, computed per layer and then averaged. The definition is simple enough to audit:

def top1_stability(previous, current):
    layer_scores = []
    for before, after in zip(previous, current):
        unchanged = sum(a == b for a, b in zip(before, after))
        layer_scores.append(unchanged / len(before))
    return sum(layer_scores) / len(layer_scores)
Fig 3. Paper-reported approximate milestones for the 80M, 64-expert analysis: chance is 1/64, stability reaches 60% at step 400 (10% of training), and exceeds 95% at the end. Bars encode those reported thresholds; they are not synthetic measurements.

Stable routing supports the “same small shard, again and again” story, but it is not a causal certificate. The paper itself reports that router jitter does not improve performance, while dropout and output masking do. Dropout sharply reduces overfitting and expert knockout cost without restoring substantial router plasticity. The likely failure sits in what fixed experts learn, not simply in the fact that routes become fixed.

ComparisonPaper-reported resultScope
80M denseminimal degradation through 8×held-out CE
80M-active MoEnoticeable degradation at 4×held-out CE
Dense versus MoEranking reverses at 32×matched active parameters
Strong maskingMoE still beats dense beyond 64×regularized runs
No interventionrecovers all-unique performancenone tested
Reported qualitative breakpoints from the paper. These summarize held-out results; they are not values generated by the accounting toy.
04

The regularizer points away from the router

InterventionRepetition effectWhat it perturbs
Residual dropoutlarge mitigationsub-layer outputs
FFN / expert maskinglarge mitigationparameter outputs
Weight decaywithin noiseweight magnitude
Gradient clippingwithin varianceupdate norm
Router jitterno clear impactrouting inputs
Qualitative outcomes from the paper's regularization sweeps; multiple hyperparameter values were tested for each method.

Strong residual dropout, FFN output masking, expert dropout, and expert output masking all help. Weight decay, clipping, and router jitter do not show measurable gains in these runs. That pattern is consistent with breaking co-adapted expert functions. It does not prove a unique mechanism: dropout changes optimization and representation learning in many ways.

05

What to probe next

A decisive follow-up would vary repetition while holding the unique pool fixed by changing total optimization steps, then separately vary U at fixed R. Cross those axes with total parameters, router plasticity, and expert dropout. Report per-expert unique coverage—not only average load—and measure whether the same token clusters stay assigned after domain shifts.

The present evidence spans several scales and domains, but still uses one training recipe, nested data subsets from a shared permutation, fixed active-parameter matching, and cross-entropy-centered evaluation. The authors show seed sensitivity and downstream tasks, yet the numeric breakpoints remain recipe-specific. For the routing computation itself, continue with the site's Transformer walkthrough; for the expert networks and dropout, see the neural-network page.

References

  1. Jha, Atindra; Li, Margaret; Leskovec, Jure; Liang, Percy; and Zettlemoyer, Luke (2026). Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data. arXiv:2609.11917
  2. Shazeer, Noam et al. (2017). Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. ICLR 2017
  3. Muennighoff, Niklas et al. (2023). Scaling Data-Constrained Language Models. NeurIPS 2023