← Blog/blog/selective-label-noise-coverage-trap

The label-noise guarantee that weakens near a tie

01

A noise matrix is a map from truth to labels

A label-noise transition matrix stores Tᵢⱼ = P(ỹ = i | Y = j): the chance that true class j is recorded as label i. Knowing it lets a learner correct losses, audit biased labels, and build prediction sets when annotations are unreliable.

The usual anchor method searches for one example that a model believes is almost certainly class j, then reads that example's noisy-label posterior as column j. That makes the whole matrix depend on one extreme, pointwise probability estimate—the place calibration is least trustworthy.

02

The estimator is just a conditional histogram

For each class, the algorithm chooses a score threshold that maximizes accepted-set purity while keeping at least a prescribed number of examples. The site runs the TypeScript below; Python and C++ are line-for-line translations.

def estimate_column(labels, accepted, class_count):
    counts = [0] * class_count
    total = 0
    for label, keep in zip(labels, accepted):
        if keep:
            counts[label] += 1
            total += 1
    if total == 0:
        raise ValueError('accept at least one sample')
    return [count / total for count in counts]

On a six-example toy score list, requiring four accepted examples selects threshold 0.82, accepts 4, and estimates the target-class column entry as 0.75. The important control is coverage: accepting fewer points raises purity but leaves fewer counts from which to estimate the column.

03

Coverage buys variance by spending purity

Theorem 1 separates error into a bias term and a finite-sample term. In a deterministic illustration with 10,000 samples, ten classes, and diagonal margin 0.4, higher coverage steadily reduces sampling noise while admitting more off-class examples. The U-shape is computed from the theorem's formula; it is not a paper-reported experiment.

bias termsampling termtotal bound
Core-computed bias–variance illustration. Coverage values are 1%, 2%, 5%, 10%, 20%, 35%, 50%, 70%, and 90%; the illustrative best false-discovery rate is 0.01 + 0.12γ².
04

The guarantee has a condition number

The bound multiplies its bias by Cₜ = 2 / (Tⱼⱼ − max Tⱼₗ). The denominator is the gap between the correct observed label and the strongest wrong label in column j. When those probabilities nearly tie, the guarantee becomes weak even with a good selector.

C_T = 2 / diagonal margin
Core-computed condition number versus diagonal margin. At margin 0.05, the bias multiplier is 40; at margin 0.8, it is 2.5.
05

Avoiding a point estimate avoids the dimension curse

The anchor estimator inherits the pointwise posterior rate n⁻ᵅ⁄⁽²ᵅ⁺ᵈ⁾. With fixed sample size and smoothness, that rate approaches one as dimension grows—barely learning at all. The selective estimator's sampling term has no explicit dimension factor, though the classifier still has to rank examples well in practice.

pointwise posterior rateparametric sampling rate
Core-computed rates at n = 10,000 and smoothness α = 1. Dimensions are 2, 5, 10, 25, 50, 100, 250, and 500. Constants and selector approximation error are omitted.
06

The paper wins broadly—but not every cell

Under 0.2-uniform synthetic noise, at least one proposed algorithm improves on the anchor baseline for most model–dataset pairs. On CIFAR-10 with a neural net, Algorithm 2 reduces transition-matrix MAE from 0.009 to 0.006.

Paper-reported CIFAR-10 neural-network MAE under 0.2-uniform noise. Lower is better.
Dataset · classifierAnchorAlgorithm 1Algorithm 2
CIFAR-10 · neural net.009 ± .001.008 ± .001.006 ± .001
MNIST · neural net.016 ± .001.007 ± .000.012 ± .001
Letter · random forest.047 ± .000.021 ± .001.029 ± .001
Satellite · logistic.067 ± .023.068 ± .010.085 ± .003
Paper Table 1: transition-matrix MAE under 0.2-uniform synthetic label noise. Lower is better.

The harsh 0.45-flip setting reveals the useful caveat. On CIFAR-10 with a neural net, Algorithm 1 is worse than the anchor baseline (0.036 versus 0.029), even though Algorithm 2 reaches 0.015. The framing is powerful; the particular selector and threshold rule still matter.

Paper-reported CIFAR-10 neural-network MAE under 0.45-flip noise. The proposed Algorithm 1 loses to the anchor baseline in this configuration; Algorithm 2 wins.
07

What to probe next

The experiments inject class-dependent noise synthetically and train one binary problem per class. The paper notes that this scales poorly with class count and can struggle for minority classes. The next test should use real annotator disagreements, report per-class coverage and condition numbers, and compare downstream corrected-model accuracy—not only matrix MAE.

Build the one-versus-rest score on the logistic-regression page, or replace it with a random forest. The estimator only needs a useful ranking; it does not require calibrated probabilities.

References

  1. Xabier de Juan, Santiago Mazuelas, Yilun Zhu, Clayton Scott (2026). Estimation of the Label-Noise Transition Matrix with Performance Guarantees via Selective Classification. NeurIPS 2026; arXiv:2609.39829
  2. Giorgio Patrini et al. (2017). Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach. CVPR 2017