← Blog/blog/smooth-net-benefit-warm-start

The clinical loss that still needs likelihood first

01

A treatment threshold is a utility statement

A clinical risk model often ends in a yes-or-no action: treat when predicted risk exceeds a threshold. That threshold is not merely a classifier setting. It encodes how many unnecessary treatments are worth one correctly treated case. At a 25% threshold, the false-positive weight is t/(1 − t) = 0.333: one true positive offsets three false positives.

Net Benefit (NB) turns that trade-off into a per-patient score: true-positive treatments add one, false-positive treatments subtract the threshold odds, and untreated cases add zero. Gorgels and colleagues ask a natural question: why train for probability likelihood everywhere if the deployed decision only cares about crossing a particular threshold?

import math

def net_benefit(probabilities, labels, threshold):
    odds = threshold / (1 - threshold)
    utility = 0.0
    for probability, label in zip(probabilities, labels):
        if probability > threshold:
            utility += 1.0 if label == 1 else -odds
    return utility / len(probabilities)

def smooth_net_benefit(probabilities, labels, threshold, k):
    odds = threshold / (1 - threshold)
    utility = 0.0
    for probability, label in zip(probabilities, labels):
        treatment = 1 / (1 + math.exp(-k * (probability - threshold)))
        utility += treatment * (1.0 if label == 1 else -odds)
    return utility / len(probabilities)

The charts run the TypeScript implementation above; Python and C++ are line-for-line translations. All examples below are deterministic accounting exercises unless explicitly labeled as paper-reported results.

02

Smooth the decision, and a gradient appears

Hard treatment is an indicator: it jumps from zero to one at the threshold, so its gradient is zero almost everywhere. Smooth Net Benefit (σNB) replaces that jump with a sigmoid. The inverse temperature k controls the bargain: low k spreads gradient widely but poorly approximates the decision; high k resembles the hard rule but concentrates learning into a narrow band.

k = 1k = 4k = 10
Fig 1. Smooth treatment weight around a 25% decision threshold. The paper anneals k from 1 to 4 to 10. Values are computed from the published sigmoid definition.
k = 1k = 4k = 10
Fig 2. Absolute positive-case gradient before model-chain factors. A sharper surrogate gives a taller but narrower learning signal, exposing the approximation-versus-optimization trade-off.
03

Even the stopping rule switches objectives

Increasing smooth NB does not guarantee increasing hard NB. The authors therefore inspect checkpoints with the actual thresholded metric. For logistic regression and GAM, that selection uses training-set NB; XGBoost uses inner validation folds. Only the final comparison uses untouched outer test folds.

StageOptimized or measuredData / dependency
1. InitializeBernoulli NLLcross-validated NLL and regularization
2. ContinueσNB, with k = 1 → 4 → 10retain NLL-selected penalties for LR/GAM
3. Stophard NBtraining NB for LR/GAM; inner-fold NB for XGBoost
4. Evaluatehard NB on untouched dataheld-out outer test folds
The paper's full recipe. The headline objective is only the middle stage.

A four-patient toy makes the mismatch visible. The smooth score changes as probabilities move, while hard NB stays on plateaus and then drops whenever a negative case crosses 50%. Our selector keeps step 0, the earliest checkpoint with hard NB 0.500.

Toy stepNLL ↓σNB, k=4 ↑hard NB ↑
00.5110.0990.500
10.4590.1310.500
20.4140.1590.250
30.3940.1750.250
Seed-free illustrative predictions for labels [1, 0, 1, 0]. These are not paper measurements.
04

The benchmark mostly favors model flexibility

Across the paper's restricted TabZilla analysis, σNB lifts mean standardized NB for logistic regression by 0.0096, with a 95% confidence interval from −0.0001 to 0.0193. It lowers the mean for GAM and does not beat NLL for any XGBoost Hessian implementation. In the reduced-data analysis, the logistic advantage reverses to −0.0037. The striking result is not that the new loss wins; it is that XGBoost's model-class advantage is much larger than the objective swap.

Fig 3. Paper-reported mean standardized-NB difference, σNB minus NLL, over 59 retained dataset-threshold combinations. The XGBoost bar uses the best reported σNB Hessian result. Differences are computed from Table 2.
Model familyNLLbest σNBσNB − NLL
Logistic regression0.56690.57650.0096
GAM0.59210.5625-0.0296
XGBoost (best σNB)0.67450.6735-0.0010
Paper-reported mean standardized Net Benefit in the restricted TabZilla analysis.
05

What the result does—and does not—license

The main analysis began with 72 dataset-threshold combinations and retained 59 after removing cases where both logistic models were near the floor or ceiling; unfiltered results are supplied in the appendix. Threshold bands were fixed at ±0.025 around prevalence-derived centers, and there was no dedicated benchmark of large clinical datasets. Those choices are reasonable for exploration, but they narrow the claim.

The most informative next experiment is a factorial ablation: random versus NLL initialization, objective-specific versus NLL-selected regularization, and smooth-surrogate versus hard-NB checkpoint selection. That would separate the value of σNB itself from the likelihood scaffold that currently makes it train. You can inspect the underlying probability model on the logistic-regression page, then compare the model-class alternative on the XGBoost page. The illustrative threshold-range calculation in this post evaluates to 0.443.

References

  1. Koen M.F. Gorgels, Lasai Barreñada, Maarten van Smeden, Ben Van Calster, Ewout W. Steyerberg, and Wouter A.C. van Amsterdam (2026). Optimizing for the decision not the prediction: an exploration of Smooth Net Benefit as a training objective. arXiv preprint
  2. Andrew J. Vickers and Elena B. Elkin (2006). Decision curve analysis: A novel method for evaluating prediction models. Medical Decision Making 26(6)
  3. Tilmann Gneiting and Adrian E. Raftery (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association 102(477)