A model leaderboard is also a calendar
Daily PM10 is the concentration of airborne particles smaller than ten micrometres. Monahan and da Silva forecast it one to seven days ahead at 425 European background stations and 365 U.S. monitors. Their primary candidates are persistence, SARIMA, and XGBoost: yesterday's value, a classical seasonal time-series model, and a nonlinear boosted-tree model.
The experiment changes the evaluation protocol while holding stations, forecast horizons, and candidates fixed. A static split fits once and scores a later block. Rolling-origin evaluation repeatedly advances time and refits, closer to how a deployed forecaster would be updated. Neither is a clerical detail: changing between them selects a different winner at 35.3% of European stations and 31.5% of U.S. monitors.
Measure movement, not just the winning name
At each horizon, the paper orders models from lowest to highest loss and compares two complete rankings with Kendall's tau. Tau is 1 when every pair agrees, −1 under complete reversal, and values in between when only some pairs invert. The Protocol Sensitivity Score (PSS) averages 1 − tau across horizons, so it runs from 0 to 2.
def kendall_tau(a, b):
pos = {model: i for i, model in enumerate(b)}
concordant, discordant = 0, 0
for i in range(len(a)):
for j in range(i + 1, len(a)):
if pos[a[i]] < pos[a[j]]:
concordant += 1
else:
discordant += 1
return (concordant - discordant) / (concordant + discordant)
def pss(rankings_a, rankings_b, weights):
return sum(
w * (1 - kendall_tau(a, b))
for a, b, w in zip(rankings_a, rankings_b, weights)
)The site runs the TypeScript implementation; Python and C++ are faithful translations. With three models and seven equally weighted horizons, one pair inversion changes the average PSS by 0.095. With nine models it changes by only 0.0079. That resolution change matters when comparing candidate-set sizes.
Protocol choice dominates the chosen conventions
| Panel | Measure | Between protocols | Rolling convention | Static convention |
|---|---|---|---|---|
| Europe (EEA) | PSS | 0.801 [0.660, 0.908] | 0.230 [0.212, 0.257] | 0.072 [0.064, 0.083] |
| Europe (EEA) | Winner swap | 35.3% [22.1%, 47.6%] | 4.9% [3.0%, 8.1%] | 0.8% [0.4%, 1.3%] |
| United States (EPA) | PSS | 0.772 [0.711, 0.823] | 0.237 [0.197, 0.277] | 0.065 [0.044, 0.085] |
| United States (EPA) | Winner swap | 31.5% [20.6%, 38.9%] | 6.3% [3.4%, 10.0%] | 1.2% [0.4%, 2.1%] |
The spatial bootstrap is important. Nearby monitors share weather and emissions, so 425 rows are not 425 independent experiments. The paper blocks European stations by country and checks alternative geographic blocks; the interval ordering remains intact.
Target pairing changes the scientific conclusion
This is the paper's most reusable lesson. If two validation schemes score different dates, a ranking can move because those dates are simply easier for a different model. Pairing turns “same task” from a verbal promise into a checkable set equality over date–horizon cells.
More candidates destabilize the top even as PSS falls
Expanding from three to nine models raises the European winner-swap rate from 35.3% to 60.2%, but mean PSS falls from 0.801 to 0.540. There is no contradiction. Many lower-ranked pairwise relations remain stable while close contenders trade places at the top. Kendall's tau weights every pair; deployment cares disproportionately about rank one.
What to probe next
Repeat the audit on classification folds, covariate-shift benchmarks, and foundation-model harnesses; pre-register tie handling; preserve example identity across every contrast; and report both full-ranking displacement and Top-1 swaps. For time series, also vary the amount and recency of training data separately from the refit schedule so those mechanisms do not travel together.
The primary nonlinear candidate is explained on the XGBoost model page. For a complementary case where point forecasts depend on what information existed at evaluation time, see the forecast-provenance walkthrough.
References
- K. Monahan and R. da Silva (2026). Quantifying Protocol-Induced Uncertainty in Comparative Predictive-Model Evaluation: Evidence from Large-Scale Daily PM10 Forecasting. arXiv:2609.26288 [cs.LG]
- M. G. Kendall (1938). A New Measure of Rank Correlation. Biometrika 30(1–2), 81–93
- L. J. Tashman (2000). Out-of-sample tests of forecasting accuracy: an analysis and review. International Journal of Forecasting 16(4), 437–450