← Blog/blog/long-tail-oversampling-reuse-tax

Five images sampled equally are still five images

01

Equal exposure is not equal information

A long-tailed dataset has a few common classes and many rare ones. In the paper's hardest CIFAR-100-LT split, the largest class has 500 training images and the smallest has five. Drawing examples uniformly from the 10,847-image dataset makes the common class appear about 5.90times in a 128-image batch. The rare class appears only 0.059times in expectation and is absent from 94.3% of batches.

Class-balanced sampling seems like the obvious repair: choose one of the 100 labels uniformly, then choose an image inside that label. But the rare label still contains only five distinct images. Its greater batch presence is repeated exposure, not new evidence.

02

Four samplers redistribute the same finite evidence

Instance-uniform sampling gives class k probability nₖ/n. Class-balanced sampling gives every class 1/K. Square-root sampling is the compromise √nₖ/Σ√nⱼ. The progressive rule linearly interpolates from instance-uniform to class-balanced over 200 epochs. Sampling is with replacement, and each epoch draws one dataset-length worth of examples.

import math

def class_probabilities(counts, strategy, epoch=0, total_epochs=200):
    total = sum(counts)
    uniform = [count / total for count in counts]
    balanced = [1 / len(counts) for _ in counts]
    if strategy == 'uniform':
        return uniform
    if strategy == 'balanced':
        return balanced
    if strategy == 'sqrt':
        denominator = sum(math.sqrt(count) for count in counts)
        return [math.sqrt(count) / denominator for count in counts]
    lam = epoch / (total_epochs - 1)
    return [(1 - lam) * u + lam * b for u, b in zip(uniform, balanced)]

The site executes the stricter TypeScript module in core/; the three snippets express the same four probability rules line for line. The reconstruction is deterministic and uses no ML library.

uniform tailsquare-root tailprogressive tailclass-balanced tail
Fig 1. Probability of drawing the five-image tail class over the paper's 200 epochs. Uniform, square-root, and balanced rules are fixed; progressive sampling ends at 1%. Computed from the published class-count and schedule formulas.
03

One hundred draws reveal only five pictures

With replacement, m draws from a class containing nₖ images expose an expected nₖ[1 − (1 − 1/nₖ)ᵐ] distinct images. The first few rare-class draws help. Then the curve saturates: after 100 draws, essentially all five images have appeared, but each unique image has been reused about 20.0times on average. A 500-image head class is still revealing new images.

unique of 5 tail imagesunique of 500 head images
Fig 2. Expected distinct examples after repeated within-class draws. The tail curve saturates at five; the head curve remains nearly linear over this range. Exact occupancy expectation, illustrative class sizes from the paper—not a training run.

The paper formalizes a related noise cost. Under independent with-replacement batches and a common within-class gradient variance, class balancing multiplies the rare class's gradient-estimator variance relative to uniform sampling by n²/(K²nₖ²). For n = 10,847, K = 100, and nₖ = 5, the multiplier is 470.6×.

04

The reported reversal is group-wide

The table and chart below reproduce the paper's 100:1 results for a ResNet-32 trained with cross-entropy. “Head,” “medium,” and “tail” group classes by their training counts. Every class-balanced group scores below its uniform counterpart. Progressive sampling trades 1.8 head points for 2.7 tail points and lands almost level overall.

SamplerOverall (%)Head (%)Medium (%)Tail (%)
Uniform39.766.738.410.8
Class-balanced3457.732.19.2
Square-root38.163.836.510.9
Progressive4064.938.613.5
Paper-reported mean test accuracy over seeds 42, 123, and 456 at 100:1 imbalance.
Fig 3. Paper-reported group accuracy at 100:1 imbalance. The highlighted bar is progressive sampling's tail result. Values are transcribed from Table 6; no synthetic observations are mixed in.
05

A curriculum—or a learning-rate interaction?

Progressive sampling looks like a curriculum: learn broad visual features from the larger classes, then emphasize rare labels. Yet the experiment changes another control at almost the same moment. At epoch 160, λ has reached 0.804 and the learning rate drops 100×; at epoch 180, it drops another 100×. The paper does not shift or remove those milestones, so the sampler schedule and optimizer schedule are not independently identified.

balanced share λlearning rate / max
Fig 4. Progressive balanced share and learning rate as fractions of their maxima. The rare-class emphasis becomes strongest just as optimization steps collapse. Both curves are reconstructed from the paper's protocol.
ImbalanceSeed 1 Δ overallSeed 2 Δ overallSeed 3 Δ overall
10:11.420.750.77
50:10.540.42-0.85
100:10.610.52-0.27
Paper-reported paired differences in percentage points: progressive minus uniform. Negative means uniform won that seed.
06

What to probe next

A clean follow-up is a small factorial experiment: hold the sampler schedule fixed while moving the learning-rate drops, then hold the optimizer fixed while changing when λ rises. Select checkpoints on a validation split, evaluate the test set once, and report many more seeds. Track per-class training loss and unique-example reuse alongside accuracy; they can distinguish useful late emphasis from memorization.

The practical lesson extends beyond vision. Reweighting a loss or resampling a label can correct exposure, but it cannot manufacture diversity. Start with the neural-network walkthroughfor the optimization machinery, then compare the related hard-example behavior in the focal-loss walkthrough.

References

  1. Siyu Yuan (2026). Mini-batch Sampling Strategies for Long-Tailed Image Classification: An Empirical Study on CIFAR-100-LT. arXiv preprint
  2. Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song, and Serge Belongie (2019). Class-Balanced Loss Based on Effective Number of Samples. CVPR 2019