The cheapest label is the one you never request
Suppose an incumbent classifier is already serving traffic and a candidate is waiting to replace it. Run both on the same input. If their predicted labels agree, then under zero-one loss they are either both right or both wrong. Their loss difference is exactly zero, even before the true label arrives. Only disagreements can change which model wins.
Balachandran's paper turns that elementary fact into Discern, a sequential audit for model updates. A sequential auditcan be checked continuously without invalidating its error guarantee. It first watches disagreements without buying labels; if that cannot settle the decision, it labels only a sampled fraction of disagreements.
The risk difference lives on disagreement
Let D be candidate loss minus incumbent loss. Its expectation Δ is the update's risk difference: Δ > 0 means the candidate is worse. Let A equal one when the predictions disagree, and let ρ be the probability of disagreement. For any bounded loss determined by the prediction, agreement forces D = 0, so Δ = E[DA] and |Δ| ≤ Bρ, where B bounds the absolute loss difference. For zero-one loss, B = 1 and Δ is the accuracy drop.
The six-row deterministic example used by this page has disagreement rate 66.7%. Its full paired difference is 0.000, and multiplying every contribution by A gives the same 0.000. That equality, not a learned surrogate, is the engine.
With a 0.9% upper bound on ρ, Tier 0 fires at a 1% tolerance: true. The strict inequality matters. An upper bound of exactly 1% does not certify the strict claim Δ < 1%.
Sample disagreements, then repay the sampling debt
In Tier 1, a label is requested on a disagreement with a probability π fixed before its outcome is seen. A queried loss difference is divided by π. This is a Horvitz–Thompson correction: rare queries receive larger weight so the average remains unbiased. With π = 0.25 and D = 1, the observed increment is 4 one quarter of the time and zero otherwise, giving expectation 1.
def audit_increment(old_pred, new_pred, label, queried, query_prob):
old_loss = int(old_pred != label)
new_loss = int(new_pred != label)
difference = new_loss - old_loss
disagrees = int(old_pred != new_pred)
if not disagrees or not queried:
return 0.0
return difference / query_prob
def verdict_claims(lower, upper, tolerance):
return {
'safe': upper < tolerance,
'regression': lower > 0,
}The snippets express the same increment and the same two verdict claims line for line. The page executes the pure TypeScript module in core/. At 10,000 stream points, 8% disagreement, and a 25% query rate, the expected ledger is 200 labels.
The words safe and regression overlap
Here is the buried semantic hinge. Discern's Safe(ε) verdict asserts Δ < ε. Its Regression verdict asserts Δ > 0. These are not logical opposites. If a confidence interval lies inside (0, ε), both statements are true: the candidate is worse, but by less than the accepted tolerance.
Noninferiority means ruling out harm larger than a prespecified margin, not proving improvement. That can be the right product decision: a tiny accuracy cost may buy lower latency or memory. But the margin is a policy choice with units. Calling it merely “safe” hides the trade.
Pairing buys labels by spending compute
In the audited regime, the leading pairing-aware label rate scales like ρ²/ε², while an auditor that labels traffic without observing both predictions pays ρ/ε². Ignoring logarithmic and range terms, the ratio is 1/ρ. The saving is not free: both models must shadow-score every audited input.
| Method | Tier-0 band labels | Audited-band labels |
|---|---|---|
| Uniform labeling | 1,476 | 2,098 |
| PPI-style | 2,232 | 2,845 |
| Active testing | 1,474 | 1,948.5 |
| Discern | 0 | 70 |
The guarantee has a stationary address
The primary theorem assumes a stationary independent stream. The paper deliberately tests what happens when the target moves: the unrestarted 95% confidence sequence miscovers on 6.4% of smooth-drift streams and 15.9% of adversarial-burst streams. Those rows are outside the theorem, not contradictions of it. A 2,000-point restarted window reduces the corresponding measurements to 0.01% and 6.5%, trading efficiency for freshness.
| Campaign | Streams | Miscoverage | Power |
|---|---|---|---|
| Pre-registered pilot | 2,352 | 0.21% | 100.0% |
| Main battery | 4,290 | 0.02% | 98.6% |
| Held-out confirmation | 2,145 | 0.09% | 97.0% |
| LLMs: 160M–410M | 80 | 2.50% | 92.3% |
| LLMs: 1B–1.4B | 60 | 1.67% | 91.7% |
What to probe next
First, write the tolerance in product units before looking at data: how many accuracy points may latency, cost, or fairness gains purchase? Report “within 1%” rather than “no regression.” Second, log the query probability before any label or complaint arrives, and audit that order. Third, run both an all-history confidence sequence for the traffic-average claim and a windowed monitor for regime-local failures. Finally, predefine slices; a good global average does not protect an unknown subgroup.
The model itself can be as simple as the classifier in the logistic-regression walkthrough. The new idea lives in paired evaluation and stopping, not in the training algorithm. For larger language models, continue with the transformer walkthrough, while keeping the paper's classification-only scope in view.
References
- Vishnu Bindu Balachandran (2026). Pay Only for Disagreement: Certified No-Regression Verdicts for Model Updates with Matching Label-Complexity Bounds. arXiv preprint
- Steven R. Howard, Aaditya Ramdas, Jon McAuliffe, and Jasjeet Sekhon (2021). Time-uniform, nonparametric, nonasymptotic confidence sequences. The Annals of Statistics