The useful part ignores identity
Chronos-2 can place several time series in one group and let them attend to each other. In a multivariate group, the series are different sensors observed over the same time window. In an in-context learning (ICL) group, they are example windows from the same or similar sensor at other times. The same attention mechanism serves both jobs.
Fore and colleagues split each cross-series attention head into two pathways. The value/output pathway, V/O, projects a pooled summary. The query/key pathway, Q/K, decides which series receives more weight. Their inference-time edits show that pooling is broadly useful, while learned identity-dependent weighting often subtracts value from ICL examples.
One attention row is just a weighted average
At a fixed patch and head, each row of the attention matrix is a softmax distribution over the V series in the group. For target series i, the update is Σⱼ αᵢⱼWᴼWⱽxⱼ. Replacing α with the identity matrix blocks cross-series flow; replacing it with 1/V keeps the learned V/O projection but erases Q/K routing.
def group_update(values, scores, mode):
if mode == "identity":
weights = [1.0] + [0.0] * (len(values) - 1)
elif mode == "uniform":
weights = [1.0 / len(values)] * len(values)
else:
peak = max(scores)
exp_scores = [exp(score - peak) for score in scores]
total = sum(exp_scores)
weights = [value / total for value in exp_scores]
return sum(weight * value for weight, value in zip(weights, values))For a deterministic three-value example [2, 5, 11], logits proportional to [1, 2, 3] produce weights 0.167, 0.333, 0.500. The learned pool is 7.5, while the uniform pool is 6.0. These are illustrative core calculations, not paper measurements.
A progressive ladder separates pooling from routing
The experiment restores one component at a time: identity α reproduces univariate inference; uniform α adds pooling; row-shuffled α preserves each row's entropy and self-weight but scrambles which series receives each off-diagonal weight; learned α restores the shipped model. Scores are percentage points of mean absolute error improvement over the univariate baseline, so the three contributions add exactly.
The sign flips with the grouping regime
A multivariate group is naturally aligned by clock time: the same storm, load peak, or tide phase reaches related sensors. An ICL group contains disjoint historical windows, so equal patch indices need not describe the same phase. The authors hypothesize that this missing alignment explains why Q/K routing helps multivariate groups but hurts ICL groups.
| ICL group | Context | Pooling | Shuffled | Learned | Total |
|---|---|---|---|---|---|
| Atlantic tides | 204 | 7.59 | 2 | -1.85 | 7.74 |
| Atlantic tides | 320 | 10.89 | 2.72 | -3.81 | 9.8 |
| Pacific tides | 204 | 13.77 | 4.14 | -16.04 | 1.87 |
| Pacific tides | 320 | 14.1 | 4.01 | -14.48 | 3.63 |
| USCRN | 204 | 1.71 | 0.3 | 0.14 | 2.15 |
| USCRN | 320 | 1.03 | -0.02 | -0.46 | 0.55 |
| PJM load | 204 | -0.87 | 0.42 | -3.32 | -3.77 |
| PJM load | 320 | 0.41 | 0.18 | -7.52 | -6.94 |
| EU load | 204 | -0.05 | 1.19 | -8.17 | -7.02 |
| EU load | 320 | 3.2 | 0 | -18.59 | -15.39 |
Depth compounds harm, but block 0 is the cheap break
Turning on one group-attention block at a time finds only block 0 has a consequential isolated effect; no other single block moves the forecast by more than 0.52 points. Activating prefixes of the 12-block stack, however, drives the worst ICL contribution to −16.09 points. The harm is therefore a property of composition across depth, not one obviously bad late block.
Starting from the shipped model and forcing block 0 to uniform improves every ICL configuration, materially in 8 of 10, by 0.4 to 20.5 points. Uniforming every block removes 11.8% of model parameters but improves ICL by 5.92 points on average; changing block 0 alone does better at 7.2.
What the experiment does—and does not—support
The clean lesson is operational: when example windows are not aligned, compare learned routing with a uniform-attention null and keep the univariate baseline. A learned attention map is not automatically more informative than a fixed pool. The broader claim “attention weights are useless” would be wrong: learned weighting helps 8 of 10 multivariate sensor-network configurations, and the learned V/O projection beats a plain cross-series mean in the paper's control.
Next, probe phase alignment directly, repeat the intervention on other architectures, and retrain the modified design rather than only patching a checkpoint. For the underlying mechanism, see the interactive Transformer walkthrough; the softmax and weighted-average pieces above are the same primitives, applied across series rather than tokens.
Reproduction notes
The charts labeled paper-reported transcribe Table A1 and the main-text counts. The temperature and three-value examples are deterministic, illustrative calculations from the post's pure core functions. No external assets, stochastic simulation, or ML library is used.
References
- Fore, M.; Inder, J. M.; Nair, M.; Vaddamanu, P.; Keshava, S. (2026). Pooling Helps, Learned Weighting Hurts In-Context: Decomposing Group Attention. arXiv:2610.01831
- Ansari, A. F. et al. (2025). Chronos-2: From Univariate to Universal Forecasting. arXiv:2510.15821