← Blog/blog/disagreement-audit-tolerance-gap

“Safe” can still mean a 0.99-point accuracy drop

01

The cheapest label is the one you never request

Suppose an incumbent classifier is already serving traffic and a candidate is waiting to replace it. Run both on the same input. If their predicted labels agree, then under zero-one loss they are either both right or both wrong. Their loss difference is exactly zero, even before the true label arrives. Only disagreements can change which model wins.

Balachandran's paper turns that elementary fact into Discern, a sequential audit for model updates. A sequential auditcan be checked continuously without invalidating its error guarantee. It first watches disagreements without buying labels; if that cannot settle the decision, it labels only a sampled fraction of disagreements.

02

The risk difference lives on disagreement

Let D be candidate loss minus incumbent loss. Its expectation Δ is the update's risk difference: Δ > 0 means the candidate is worse. Let A equal one when the predictions disagree, and let ρ be the probability of disagreement. For any bounded loss determined by the prediction, agreement forces D = 0, so Δ = E[DA] and |Δ| ≤ Bρ, where B bounds the absolute loss difference. For zero-one loss, B = 1 and Δ is the accuracy drop.

The six-row deterministic example used by this page has disagreement rate 66.7%. Its full paired difference is 0.000, and multiplying every contribution by A gives the same 0.000. That equality, not a learned surrogate, is the engine.

worst possible accuracy dropSafe tolerance
Fig 1. The worst possible zero-one-loss gap Bρ against the paper's primary 1% tolerance. Rates are exact values from the support bound, not experimental observations. A confidence-sequence upper bound on ρ must be strictly below 1% for the zero-label certificate to fire.

With a 0.9% upper bound on ρ, Tier 0 fires at a 1% tolerance: true. The strict inequality matters. An upper bound of exactly 1% does not certify the strict claim Δ < 1%.

03

Sample disagreements, then repay the sampling debt

In Tier 1, a label is requested on a disagreement with a probability π fixed before its outcome is seen. A queried loss difference is divided by π. This is a Horvitz–Thompson correction: rare queries receive larger weight so the average remains unbiased. With π = 0.25 and D = 1, the observed increment is 4 one quarter of the time and zero otherwise, giving expectation 1.

def audit_increment(old_pred, new_pred, label, queried, query_prob):
    old_loss = int(old_pred != label)
    new_loss = int(new_pred != label)
    difference = new_loss - old_loss
    disagrees = int(old_pred != new_pred)
    if not disagrees or not queried:
        return 0.0
    return difference / query_prob

def verdict_claims(lower, upper, tolerance):
    return {
        'safe': upper < tolerance,
        'regression': lower > 0,
    }

The snippets express the same increment and the same two verdict claims line for line. The page executes the pure TypeScript module in core/. At 10,000 stream points, 8% disagreement, and a 25% query rate, the expected ledger is 200 labels.

04

The words safe and regression overlap

Here is the buried semantic hinge. Discern's Safe(ε) verdict asserts Δ < ε. Its Regression verdict asserts Δ > 0. These are not logical opposites. If a confidence interval lies inside (0, ε), both statements are true: the candidate is worse, but by less than the accepted tolerance.

Noninferiority means ruling out harm larger than a prespecified margin, not proving improvement. That can be the right product decision: a tiny accuracy cost may buy lower latency or memory. But the margin is a policy choice with units. Calling it merely “safe” hides the trade.

05

Pairing buys labels by spending compute

In the audited regime, the leading pairing-aware label rate scales like ρ²/ε², while an auditor that labels traffic without observing both predictions pays ρ/ε². Ignoring logarithmic and range terms, the ratio is 1/ρ. The saving is not free: both models must shadow-score every audited input.

pairing-aware: ρ²/ε²pairing-blind: ρ/ε²
Fig 2. Leading theoretical label-complexity terms at ε = 1% over disagreement rates 2%–20%. These are normalized rates from the paper's bounds, not predicted absolute label counts; logarithmic, constant, and lower-order terms are omitted.
MethodTier-0 band labelsAudited-band labels
Uniform labeling1,4762,098
PPI-style2,2322,845
Active testing1,4741,948.5
Discern070
Paper-reported median labels to a sound verdict at ε = 1% and δ = 0.05. Tier-0 n = 246; audited band n = 29. Discern resolved 28 of 29 audited pairs.
Fig 3. Paper-reported audited-band medians from the dedicated baseline comparison. Discern used 70 labels versus 2,098 for uniform labeling. Absolute values are transcribed from Table 3.
06

The guarantee has a stationary address

The primary theorem assumes a stationary independent stream. The paper deliberately tests what happens when the target moves: the unrestarted 95% confidence sequence miscovers on 6.4% of smooth-drift streams and 15.9% of adversarial-burst streams. Those rows are outside the theorem, not contradictions of it. A 2,000-point restarted window reduces the corresponding measurements to 0.01% and 6.5%, trading efficiency for freshness.

Fig 4. Paper-reported time-uniform miscoverage. The i.i.d. main battery is far below the 5% nominal budget; unrestarted monitoring of a moving target exceeds it under both synthetic drift schedules.
CampaignStreamsMiscoveragePower
Pre-registered pilot2,3520.21%100.0%
Main battery4,2900.02%98.6%
Held-out confirmation2,1450.09%97.0%
LLMs: 160M–410M802.50%92.3%
LLMs: 1B–1.4B601.67%91.7%
Paper-reported validity and power. Power is the share of injected regressions with Δ ≥ 2% alarmed within 5,000 stream points.
07

What to probe next

First, write the tolerance in product units before looking at data: how many accuracy points may latency, cost, or fairness gains purchase? Report “within 1%” rather than “no regression.” Second, log the query probability before any label or complaint arrives, and audit that order. Third, run both an all-history confidence sequence for the traffic-average claim and a windowed monitor for regime-local failures. Finally, predefine slices; a good global average does not protect an unknown subgroup.

The model itself can be as simple as the classifier in the logistic-regression walkthrough. The new idea lives in paired evaluation and stopping, not in the training algorithm. For larger language models, continue with the transformer walkthrough, while keeping the paper's classification-only scope in view.

References

  1. Vishnu Bindu Balachandran (2026). Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds. arXiv preprint
  2. Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon (2021). Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics