← Blog/blog/future-trained-forecast-provenance

The future-trained forecaster that got worse

01

A backtest has two clocks

A historical forecast can use only information available when the forecast was made. With a pretrained model, that rule applies twice: once to the numerical input history and again to the data that shaped the model’s parameters. A clean input window does not rescue parameters trained on the future.

Chen and coauthors make the second clock visible with annual vintages of Chronos and TimesFM trained on expanding financial datasets. For a forecast beginning in year T, the origin-aligned point-in-time (PIT)model is trained only through T−1. Vintages T, T+1, and T+2 cross the forecast origin and are temporally exposed.

02

Direct omission does not close the leak

Direct exposure means the target market appears in the post-origin training sample. Indirect exposure means later observations from related markets or factors can still encode the target’s period. The core classifies those cases as direct and indirect. Removing one target series is therefore weaker than enforcing a cutoff on the entire training corpus.

Annual stateLeadHistorical statusComparison role
T−2−2admissiblestale
T−1−1admissiblePIT benchmark
T0exposedcrosses origin
T+1 / T+2+1 / +2exposeddeeper future
Paper design. Annual state v contains training data through year-end v and becomes admissible only from year v+1.
03

Every revision pays a quadratic toll

Let e = y − ŷPIT be the error left by the PIT forecast and D = ŷalt − ŷPIT the revision from swapping model vintages. The alternative error is e−D, so the PIT-minus-alternative squared loss is exactly 2eD−D². The first term rewards movement toward the realized return; the second charges every move, even a well-intentioned one.

In the deterministic example, y=10, the PIT forecast is 8, and the later forecast is 9. Alignment contributes 4, the movement penalty is 1, and the later model gains 3 squared-error units. Reversing the move to 7 would lose five.

def decompose(actual, pit, alternative):
    error = actual - pit
    revision = alternative - pit
    alignment = 2 * error * revision
    movement = revision * revision
    return {
        'pit_error': error,
        'revision': revision,
        'alignment': alignment,
        'movement': movement,
        'gain': alignment - movement,
    }

def normalized_gain(scale, efficiency):
    if scale < 0 or not -1 <= efficiency <= 1:
        raise ValueError('invalid geometry')
    return 2 * efficiency * scale - scale * scale

Aggregate that identity and define revision scale r and alignment efficiency κ. The normalized gain becomes 2κr−r², so a revision helps only when κ>r/2. Bigger changes demand proportionally better direction.

revision scale r=0.2revision scale r=0.4revision scale r=0.6break-even
Deterministic geometry from the core formula, not paper data. X positions are κ = −0.1, 0, 0.1, …, 0.5; values above zero improve on PIT.
04

The exposed revisions moved plenty—and aligned poorly

Table 8 turns the algebra into a diagnostic. U.S. rolling revisions were 27.3% as large as the remaining PIT error, requiring κ=0.137 to break even. Actual κ was −0.014. Fixed-state deployment required 0.174 and delivered only 0.006. International results miss by similarly wide margins.

TargetsComparisonrκr / 2κ − r/2Gain cells
U.S.PIT vs stale0.2660.1360.1330.00360%
U.S.rolling exposed0.273-0.0140.137-0.15112%
U.S.fixed state0.3480.0060.174-0.16813%
InternationalPIT vs stale0.3080.1680.1540.01453%
Internationalrolling exposed0.3120.0240.156-0.13217%
Internationalfixed state0.3730.0410.186-0.14517%
Paper-reported revision scale, alignment, and gain shares; thresholds and margins recomputed by the core functions. ‘PIT vs stale’ uses the paper’s PIT-relative direction.
Chronos TinyChronos MiniChronos SmallTimesFM 8MTimesFM 20M
Paper-reported pooled rolling effects for U.S. returns under U.S. training. Positive favors the exposed state; x positions are 1, 3, 6, and 12 months.
Model1m3m6m12m
Chronos Tiny-1.29-5.88-9.26-6.01
Chronos Mini-2.66-5.45-5.750.82
Chronos Small1.86-4.42-13.72-15.24
TimesFM 8M-1.89-6.31-13.63-16.55
TimesFM 20M-3.52-15.13-23.24-20.53
Internet Appendix Table IA.4.1. Effects are PIT MSFE minus exposed-state MSFE, normalized by historical-average MSFE; negative favors PIT.
05

Crossing the origin is not an ordinary update

The adjacent-vintage comparison holds the calendar step to one year. Moving from T−2 to T−1 adds only pre-origin training data and lowers U.S. loss by 0.72 historical-average-MSFE percentage points. Moving from T−1 to T first crosses the origin and raises loss by 5.65 points. Across 13 non-U.S. markets, the corresponding effects are +0.85 and −5.71.

Paper-reported mean matched predictive effects. Positive means the newer state lowers loss; negative means it raises loss.

The forecasts also changed economically: the first origin-crossing TimesFM 20M state flipped the one-month U.S. return sign in 27.9% of matched months. Under the paper’s constrained one-month market-timing rule, exposed-minus-PIT annualized certainty-equivalent return had a −1.77 percentage-point U.S. median and a −2.14-point median across all 65 international market–model cells. Those portfolio results exclude transaction costs.

06

The clean causal experiment is still missing

“PIT” here means training-date alignment, not that the architecture or parameters actually existed at that historical date. Cutoffs are annual, deployment is monthly, daily firm-level training becomes aggregate market forecasting, and only five variants from two model families are studied. Training content matters too: for U.S. targets the pooled median shifts from −5.94 points under U.S. training to +1.37 under global training, then back to −3.64 after factor augmentation. There is no universal monotone penalty.

The next experiment should continue-train identical checkpoints across cryptographically recorded data cutoffs, repeat seeds, and report direct and indirect exposure separately. It should also charge trading costs and test finer cutoff dates. Until then, keep two audit columns: was this model historically admissible? and did it perform better? Never let the second answer overwrite the first.

For the sequence architecture behind these foundation forecasters, continue with the Transformer walkthrough. For the historical-average benchmark and forecast-error geometry, see the linear-regression walkthrough.

References

  1. H. Chen, L. Chen, Y. Chen, D. Huang, and B. Zhang (2026). Does Training on Future Data Pay? Look-Ahead Bias in Forecasting with Pretrained Models. arXiv:2609.20554
  2. A. F. Ansari et al. (2024). Chronos: Learning the Language of Time Series. Transactions on Machine Learning Research
  3. A. Das et al. (2024). A Decoder-only Foundation Model for Time-series Forecasting. ICML