A treatment threshold is a utility statement
A clinical risk model often ends in a yes-or-no action: treat when predicted risk exceeds a threshold. That threshold is not merely a classifier setting. It encodes how many unnecessary treatments are worth one correctly treated case. At a 25% threshold, the false-positive weight is t/(1 − t) = 0.333: one true positive offsets three false positives.
Net Benefit (NB) turns that trade-off into a per-patient score: true-positive treatments add one, false-positive treatments subtract the threshold odds, and untreated cases add zero. Gorgels and colleagues ask a natural question: why train for probability likelihood everywhere if the deployed decision only cares about crossing a particular threshold?
import math
def net_benefit(probabilities, labels, threshold):
odds = threshold / (1 - threshold)
utility = 0.0
for probability, label in zip(probabilities, labels):
if probability > threshold:
utility += 1.0 if label == 1 else -odds
return utility / len(probabilities)
def smooth_net_benefit(probabilities, labels, threshold, k):
odds = threshold / (1 - threshold)
utility = 0.0
for probability, label in zip(probabilities, labels):
treatment = 1 / (1 + math.exp(-k * (probability - threshold)))
utility += treatment * (1.0 if label == 1 else -odds)
return utility / len(probabilities)The charts run the TypeScript implementation above; Python and C++ are line-for-line translations. All examples below are deterministic accounting exercises unless explicitly labeled as paper-reported results.
Smooth the decision, and a gradient appears
Hard treatment is an indicator: it jumps from zero to one at the threshold, so its gradient is zero almost everywhere. Smooth Net Benefit (σNB) replaces that jump with a sigmoid. The inverse temperature k controls the bargain: low k spreads gradient widely but poorly approximates the decision; high k resembles the hard rule but concentrates learning into a narrow band.
Even the stopping rule switches objectives
Increasing smooth NB does not guarantee increasing hard NB. The authors therefore inspect checkpoints with the actual thresholded metric. For logistic regression and GAM, that selection uses training-set NB; XGBoost uses inner validation folds. Only the final comparison uses untouched outer test folds.
| Stage | Optimized or measured | Data / dependency |
|---|---|---|
| 1. Initialize | Bernoulli NLL | cross-validated NLL and regularization |
| 2. Continue | σNB, with k = 1 → 4 → 10 | retain NLL-selected penalties for LR/GAM |
| 3. Stop | hard NB | training NB for LR/GAM; inner-fold NB for XGBoost |
| 4. Evaluate | hard NB on untouched data | held-out outer test folds |
A four-patient toy makes the mismatch visible. The smooth score changes as probabilities move, while hard NB stays on plateaus and then drops whenever a negative case crosses 50%. Our selector keeps step 0, the earliest checkpoint with hard NB 0.500.
| Toy step | NLL ↓ | σNB, k=4 ↑ | hard NB ↑ |
|---|---|---|---|
| 0 | 0.511 | 0.099 | 0.500 |
| 1 | 0.459 | 0.131 | 0.500 |
| 2 | 0.414 | 0.159 | 0.250 |
| 3 | 0.394 | 0.175 | 0.250 |
The benchmark mostly favors model flexibility
Across the paper's restricted TabZilla analysis, σNB lifts mean standardized NB for logistic regression by 0.0096, with a 95% confidence interval from −0.0001 to 0.0193. It lowers the mean for GAM and does not beat NLL for any XGBoost Hessian implementation. In the reduced-data analysis, the logistic advantage reverses to −0.0037. The striking result is not that the new loss wins; it is that XGBoost's model-class advantage is much larger than the objective swap.
| Model family | NLL | best σNB | σNB − NLL |
|---|---|---|---|
| Logistic regression | 0.5669 | 0.5765 | 0.0096 |
| GAM | 0.5921 | 0.5625 | -0.0296 |
| XGBoost (best σNB) | 0.6745 | 0.6735 | -0.0010 |
What the result does—and does not—license
The main analysis began with 72 dataset-threshold combinations and retained 59 after removing cases where both logistic models were near the floor or ceiling; unfiltered results are supplied in the appendix. Threshold bands were fixed at ±0.025 around prevalence-derived centers, and there was no dedicated benchmark of large clinical datasets. Those choices are reasonable for exploration, but they narrow the claim.
The most informative next experiment is a factorial ablation: random versus NLL initialization, objective-specific versus NLL-selected regularization, and smooth-surrogate versus hard-NB checkpoint selection. That would separate the value of σNB itself from the likelihood scaffold that currently makes it train. You can inspect the underlying probability model on the logistic-regression page, then compare the model-class alternative on the XGBoost page. The illustrative threshold-range calculation in this post evaluates to 0.443.
References
- Koen M.F. Gorgels, Lasai Barreñada, Maarten van Smeden, Ben Van Calster, Ewout W. Steyerberg, and Wouter A.C. van Amsterdam (2026). Optimizing for the decision not the prediction: an exploration of Smooth Net Benefit as a training objective. arXiv preprint
- Andrew J. Vickers and Elena B. Elkin (2006). Decision curve analysis: A novel method for evaluating prediction models. Medical Decision Making 26(6)
- Tilmann Gneiting and Adrian E. Raftery (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association 102(477)