A backtest has two clocks
A historical forecast can use only information available when the forecast was made. With a pretrained model, that rule applies twice: once to the numerical input history and again to the data that shaped the model’s parameters. A clean input window does not rescue parameters trained on the future.
Chen and coauthors make the second clock visible with annual vintages of Chronos and TimesFM trained on expanding financial datasets. For a forecast beginning in year T, the origin-aligned point-in-time (PIT)model is trained only through T−1. Vintages T, T+1, and T+2 cross the forecast origin and are temporally exposed.
Direct omission does not close the leak
Direct exposure means the target market appears in the post-origin training sample. Indirect exposure means later observations from related markets or factors can still encode the target’s period. The core classifies those cases as direct and indirect. Removing one target series is therefore weaker than enforcing a cutoff on the entire training corpus.
| Annual state | Lead | Historical status | Comparison role |
|---|---|---|---|
| T−2 | −2 | admissible | stale |
| T−1 | −1 | admissible | PIT benchmark |
| T | 0 | exposed | crosses origin |
| T+1 / T+2 | +1 / +2 | exposed | deeper future |
Every revision pays a quadratic toll
Let e = y − ŷPIT be the error left by the PIT forecast and D = ŷalt − ŷPIT the revision from swapping model vintages. The alternative error is e−D, so the PIT-minus-alternative squared loss is exactly 2eD−D². The first term rewards movement toward the realized return; the second charges every move, even a well-intentioned one.
In the deterministic example, y=10, the PIT forecast is 8, and the later forecast is 9. Alignment contributes 4, the movement penalty is 1, and the later model gains 3 squared-error units. Reversing the move to 7 would lose five.
def decompose(actual, pit, alternative):
error = actual - pit
revision = alternative - pit
alignment = 2 * error * revision
movement = revision * revision
return {
'pit_error': error,
'revision': revision,
'alignment': alignment,
'movement': movement,
'gain': alignment - movement,
}
def normalized_gain(scale, efficiency):
if scale < 0 or not -1 <= efficiency <= 1:
raise ValueError('invalid geometry')
return 2 * efficiency * scale - scale * scaleAggregate that identity and define revision scale r and alignment efficiency κ. The normalized gain becomes 2κr−r², so a revision helps only when κ>r/2. Bigger changes demand proportionally better direction.
The exposed revisions moved plenty—and aligned poorly
Table 8 turns the algebra into a diagnostic. U.S. rolling revisions were 27.3% as large as the remaining PIT error, requiring κ=0.137 to break even. Actual κ was −0.014. Fixed-state deployment required 0.174 and delivered only 0.006. International results miss by similarly wide margins.
| Targets | Comparison | r | κ | r / 2 | κ − r/2 | Gain cells |
|---|---|---|---|---|---|---|
| U.S. | PIT vs stale | 0.266 | 0.136 | 0.133 | 0.003 | 60% |
| U.S. | rolling exposed | 0.273 | -0.014 | 0.137 | -0.151 | 12% |
| U.S. | fixed state | 0.348 | 0.006 | 0.174 | -0.168 | 13% |
| International | PIT vs stale | 0.308 | 0.168 | 0.154 | 0.014 | 53% |
| International | rolling exposed | 0.312 | 0.024 | 0.156 | -0.132 | 17% |
| International | fixed state | 0.373 | 0.041 | 0.186 | -0.145 | 17% |
| Model | 1m | 3m | 6m | 12m |
|---|---|---|---|---|
| Chronos Tiny | -1.29 | -5.88 | -9.26 | -6.01 |
| Chronos Mini | -2.66 | -5.45 | -5.75 | 0.82 |
| Chronos Small | 1.86 | -4.42 | -13.72 | -15.24 |
| TimesFM 8M | -1.89 | -6.31 | -13.63 | -16.55 |
| TimesFM 20M | -3.52 | -15.13 | -23.24 | -20.53 |
Crossing the origin is not an ordinary update
The adjacent-vintage comparison holds the calendar step to one year. Moving from T−2 to T−1 adds only pre-origin training data and lowers U.S. loss by 0.72 historical-average-MSFE percentage points. Moving from T−1 to T first crosses the origin and raises loss by 5.65 points. Across 13 non-U.S. markets, the corresponding effects are +0.85 and −5.71.
The forecasts also changed economically: the first origin-crossing TimesFM 20M state flipped the one-month U.S. return sign in 27.9% of matched months. Under the paper’s constrained one-month market-timing rule, exposed-minus-PIT annualized certainty-equivalent return had a −1.77 percentage-point U.S. median and a −2.14-point median across all 65 international market–model cells. Those portfolio results exclude transaction costs.
The clean causal experiment is still missing
“PIT” here means training-date alignment, not that the architecture or parameters actually existed at that historical date. Cutoffs are annual, deployment is monthly, daily firm-level training becomes aggregate market forecasting, and only five variants from two model families are studied. Training content matters too: for U.S. targets the pooled median shifts from −5.94 points under U.S. training to +1.37 under global training, then back to −3.64 after factor augmentation. There is no universal monotone penalty.
The next experiment should continue-train identical checkpoints across cryptographically recorded data cutoffs, repeat seeds, and report direct and indirect exposure separately. It should also charge trading costs and test finer cutoff dates. Until then, keep two audit columns: was this model historically admissible? and did it perform better? Never let the second answer overwrite the first.
For the sequence architecture behind these foundation forecasters, continue with the Transformer walkthrough. For the historical-average benchmark and forecast-error geometry, see the linear-regression walkthrough.
References
- H. Chen, L. Chen, Y. Chen, D. Huang, and B. Zhang (2026). Does Training on Future Data Pay? Look-Ahead Bias in Forecasting with Pretrained Models. arXiv:2609.20554
- A. F. Ansari et al. (2024). Chronos: Learning the Language of Time Series. Transactions on Machine Learning Research
- A. Das et al. (2024). A Decoder-only Foundation Model for Time-series Forecasting. ICML