A SHAP value answers a counterfactual question: how much did this feature move one prediction away from a baseline? Zhu and colleagues ask a second question: would that answer survive a few additional training examples whose labels and destinations are unknown? Their interval is a useful stress test—but only after the tree itself has been frozen.
A leaf probability is only a ratio
A classification leaf with n positive labels amongN observations predicts n / N. Putm unknown labels into that leaf and its possible output is[n / (N + m), (n + m) / (N + m)]. Small leaves widen fastest because each new observation has more leverage.
Why the greedy step is enough
With the tree, explained point, background distribution, and missing- feature rule held fixed, a SHAP value is linear in the leaf outputs. Each virtual sample earns a diminishing marginal gain, so the worst case repeatedly assigns the next sample to the leaf with the largest current gain.
def greedy_allocation(weights, covers, budget):
allocation = [0] * len(covers)
for _ in range(budget):
gains = []
for weight, cover, assigned in zip(weights, covers, allocation):
gain = weight * cover / ((cover + assigned + 1) * (cover + assigned))
gains.append(gain)
best = max(range(len(gains)), key=lambda i: gains[i])
allocation[best] += 1
correction = sum(weight * assigned / (cover + assigned)
for weight, cover, assigned in zip(weights, covers, allocation))
return allocation, correctionThe site runs the TypeScript implementation above; Python and C++ are line-for-line translations. The next chart applies the same core function to two illustrative coefficient vectors and the four leaf counts from the paper's running example.
The band does not perturb the route
Here is the assumption worth circling: the paper propagates uncertainty through leaf outputs, not through split thresholds or reach probabilities. That distinction is sensible when a reviewed tree must remain structurally stable. It is not the same as uncertainty under retraining, where a new sample can move a split and send the explained point to another leaf.
What the experiments establish
The authors compare their interval lengths with empirical min-max SHAP ranges from 100 bootstrap retrainings. At virtual-sample budgets = 2, rank correlations are moderate for single trees and high for random forests. That makes the interval a useful proxy for where explanations are fragile—even though the proxy and retraining vary different objects.
| Dataset | Model | Spearman ρ |
|---|---|---|
| Diabetes | Tree | 0.57 |
| Diabetes | Random forest | 0.83 |
| Breast cancer | Tree | 0.36 |
| Breast cancer | Random forest | 0.92 |
| Ionosphere | Tree | 0.40 |
| Ionosphere | Random forest | 0.80 |
| Mortality | Tree | 0.55 |
| Mortality | Random forest | 0.90 |
In a separate synthetic feature-ranking experiment, nominal SHAP scores 0.82 AUC, the pessimistic score 0.83, and out-of-bag SHAP 0.85. The proposed method is competitive without retraining, but it does not beat the held-out correction in that experiment.
| Layer | What the interval covers | What it leaves outside |
|---|---|---|
| Leaf outputs | Unknown labels assigned to existing leaves | New splits or changed thresholds |
| Reach probabilities | Held constant for each feature coalition | Samples taking a different path |
| Single tree | Exact greedy worst-case allocation | Exact large-ensemble allocation |
| Validation | Correlation with bootstrap interval widths | Guaranteed bootstrap coverage |
What to probe next
First, compare interval coverage—not merely rank correlation—against bootstrap retraining. Second, split the bootstrap variation into leaf-output, threshold, and topology components. Third, test whether the random-forest outer approximation becomes too conservative as trees deepen. Finally, report how often a supposedly robust top feature changes once routing is allowed to move.
References
- Chenrui Zhu, Vu-Linh Nguyen, Marie-Hélène Masson, and Sébastien Destercke (2026). Interval-valued SHAP in Tree-Based Models. ICDM 2026 / arXiv:2610.11953
- Scott M. Lundberg and Su-In Lee (2017). A Unified Approach to Interpreting Model Predictions. NeurIPS 2017
- Scott M. Lundberg, Gabriel Erion, Hugh Chen, Alex DeGrave, Jordan M. Prutkin, Bala Nair, Ronit Katz, Jonathan Himmelfarb, Nisha Bansal, and Su-In Lee (2020). From Local Explanations to Global Understanding with Explainable AI for Trees. Nature Machine Intelligence 2, 56–67