← Blog/blog/interval-shap-fixed-tree-blind-spot

The SHAP uncertainty band that keeps the tree frozen

A SHAP value answers a counterfactual question: how much did this feature move one prediction away from a baseline? Zhu and colleagues ask a second question: would that answer survive a few additional training examples whose labels and destinations are unknown? Their interval is a useful stress test—but only after the tree itself has been frozen.

01

A leaf probability is only a ratio

A classification leaf with n positive labels amongN observations predicts n / N. Putm unknown labels into that leaf and its possible output is[n / (N + m), (n + m) / (N + m)]. Small leaves widen fastest because each new observation has more leverage.

30-sample leaf50-sample leaf100-sample leaf
Illustrative interval width for three leaves as the virtual-sample budget grows. The values are computed from the paper's leaf-probability bounds; no randomness is used.
02

Why the greedy step is enough

With the tree, explained point, background distribution, and missing- feature rule held fixed, a SHAP value is linear in the leaf outputs. Each virtual sample earns a diminishing marginal gain, so the worst case repeatedly assigns the next sample to the leaf with the largest current gain.

def greedy_allocation(weights, covers, budget):
    allocation = [0] * len(covers)
    for _ in range(budget):
        gains = []
        for weight, cover, assigned in zip(weights, covers, allocation):
            gain = weight * cover / ((cover + assigned + 1) * (cover + assigned))
            gains.append(gain)
        best = max(range(len(gains)), key=lambda i: gains[i])
        allocation[best] += 1
    correction = sum(weight * assigned / (cover + assigned)
                     for weight, cover, assigned in zip(weights, covers, allocation))
    return allocation, correction

The site runs the TypeScript implementation above; Python and C++ are line-for-line translations. The next chart applies the same core function to two illustrative coefficient vectors and the four leaf counts from the paper's running example.

feature A intervalfeature B interval
Illustrative pessimistic SHAP interval widths as the total unknown-sample budget grows. Different features widen at different rates because their leaf coefficients differ.
03

The band does not perturb the route

Here is the assumption worth circling: the paper propagates uncertainty through leaf outputs, not through split thresholds or reach probabilities. That distinction is sensible when a reviewed tree must remain structurally stable. It is not the same as uncertainty under retraining, where a new sample can move a split and send the explained point to another leaf.

exact SHAP after refitting split
Illustrative one-feature stump. Moving its threshold from 0.04 to 0.06 sends x = 0.05 to the opposite leaf, flipping exact SHAP from +0.30 to −0.30 while both leaf outputs stay fixed. A leaf-only interval cannot represent this route change.
04

What the experiments establish

The authors compare their interval lengths with empirical min-max SHAP ranges from 100 bootstrap retrainings. At virtual-sample budgets = 2, rank correlations are moderate for single trees and high for random forests. That makes the interval a useful proxy for where explanations are fragile—even though the proxy and retraining vary different objects.

Paper-reported Spearman correlations between proposed SHAP interval lengths and bootstrap attribution variability. Random-forest correlations are consistently higher than single-tree correlations.
DatasetModelSpearman ρ
DiabetesTree0.57
DiabetesRandom forest0.83
Breast cancerTree0.36
Breast cancerRandom forest0.92
IonosphereTree0.40
IonosphereRandom forest0.80
MortalityTree0.55
MortalityRandom forest0.90
Paper Table V, s = 2, averaged over test instances and folds. These are paper-reported values, not outputs of the illustrative charts.

In a separate synthetic feature-ranking experiment, nominal SHAP scores 0.82 AUC, the pessimistic score 0.83, and out-of-bag SHAP 0.85. The proposed method is competitive without retraining, but it does not beat the held-out correction in that experiment.

LayerWhat the interval coversWhat it leaves outside
Leaf outputsUnknown labels assigned to existing leavesNew splits or changed thresholds
Reach probabilitiesHeld constant for each feature coalitionSamples taking a different path
Single treeExact greedy worst-case allocationExact large-ensemble allocation
ValidationCorrelation with bootstrap interval widthsGuaranteed bootstrap coverage
The useful reading is conditional: robust leaf estimates are not automatically robust routes or robust refits.
05

What to probe next

First, compare interval coverage—not merely rank correlation—against bootstrap retraining. Second, split the bootstrap variation into leaf-output, threshold, and topology components. Third, test whether the random-forest outer approximation becomes too conservative as trees deepen. Finally, report how often a supposedly robust top feature changes once routing is allowed to move.

References

  1. Chenrui Zhu, Vu-Linh Nguyen, Marie-Hélène Masson, and Sébastien Destercke (2026). Interval-valued SHAP in Tree-Based Models. ICDM 2026 / arXiv:2610.11953
  2. Scott M. Lundberg and Su-In Lee (2017). A Unified Approach to Interpreting Model Predictions. NeurIPS 2017
  3. Scott M. Lundberg, Gabriel Erion, Hugh Chen, Alex DeGrave, Jordan M. Prutkin, Bala Nair, Ronit Katz, Jonathan Himmelfarb, Nisha Bansal, and Su-In Lee (2020). From Local Explanations to Global Understanding with Explainable AI for Trees. Nature Machine Intelligence 2, 56–67