← Blog/blog/group-attention-weighting-tax

The attention head that improves when its weights are erased

01

The useful part ignores identity

Chronos-2 can place several time series in one group and let them attend to each other. In a multivariate group, the series are different sensors observed over the same time window. In an in-context learning (ICL) group, they are example windows from the same or similar sensor at other times. The same attention mechanism serves both jobs.

Fore and colleagues split each cross-series attention head into two pathways. The value/output pathway, V/O, projects a pooled summary. The query/key pathway, Q/K, decides which series receives more weight. Their inference-time edits show that pooling is broadly useful, while learned identity-dependent weighting often subtracts value from ICL examples.

02

One attention row is just a weighted average

At a fixed patch and head, each row of the attention matrix is a softmax distribution over the V series in the group. For target series i, the update is Σⱼ αᵢⱼWᴼWⱽxⱼ. Replacing α with the identity matrix blocks cross-series flow; replacing it with 1/V keeps the learned V/O projection but erases Q/K routing.

def group_update(values, scores, mode):
    if mode == "identity":
        weights = [1.0] + [0.0] * (len(values) - 1)
    elif mode == "uniform":
        weights = [1.0 / len(values)] * len(values)
    else:
        peak = max(scores)
        exp_scores = [exp(score - peak) for score in scores]
        total = sum(exp_scores)
        weights = [value / total for value in exp_scores]
    return sum(weight * value for weight, value in zip(weights, values))

For a deterministic three-value example [2, 5, 11], logits proportional to [1, 2, 3] produce weights 0.167, 0.333, 0.500. The learned pool is 7.5, while the uniform pool is 6.0. These are illustrative core calculations, not paper measurements.

effective group size
Deterministic illustration. The x positions are softmax temperatures 0.25, 0.5, 1, 2, and 4. Effective group size is exp(entropy): one for a one-hot row and V for a uniform row.
03

A progressive ladder separates pooling from routing

The experiment restores one component at a time: identity α reproduces univariate inference; uniform α adds pooling; row-shuffled α preserves each row's entropy and self-weight but scrambles which series receives each off-diagonal weight; learned α restores the shipped model. Scores are percentage points of mean absolute error improvement over the univariate baseline, so the three contributions add exactly.

cumulative improvement
Paper-reported Pacific-tides ICL result at context 204. Starting from zero, uniform pooling adds 13.77 points, shuffled weighting adds 4.14, and learned identity-dependent routing removes 16.04, leaving 1.87 points.
04

The sign flips with the grouping regime

A multivariate group is naturally aligned by clock time: the same storm, load peak, or tide phase reaches related sensors. An ICL group contains disjoint historical windows, so equal patch indices need not describe the same phase. The authors hypothesize that this missing alignment explains why Q/K routing helps multivariate groups but hurts ICL groups.

uniform poolinglearned weightingtotal shipped mechanism
Paper-reported ICL contributions for Atlantic tides (204, 320), Pacific tides (204, 320), USCRN (204, 320), PJM load (204, 320), and EU load (204, 320), in that order. Positive means lower MAE than the previous rung or univariate baseline; negative means worse.
ICL groupContextPoolingShuffledLearnedTotal
Atlantic tides2047.592-1.857.74
Atlantic tides32010.892.72-3.819.8
Pacific tides20413.774.14-16.041.87
Pacific tides32014.14.01-14.483.63
USCRN2041.710.30.142.15
USCRN3201.03-0.02-0.460.55
PJM load204-0.870.42-3.32-3.77
PJM load3200.410.18-7.52-6.94
EU load204-0.051.19-8.17-7.02
EU load3203.20-18.59-15.39
All values are paper-reported percentage points of univariate-baseline MAE. Highlighted totals mean the complete shipped group-attention path is worse than univariate inference.
Paper-reported configuration counts. Pooling is positive in 18/20 MV+ICL sensor-network configurations; learned weighting is negative in 9/10 ICL configurations; the full mechanism loses to univariate in 4/10; block-0 uniform repair helps 10/10 ICL configurations.
05

Depth compounds harm, but block 0 is the cheap break

Turning on one group-attention block at a time finds only block 0 has a consequential isolated effect; no other single block moves the forecast by more than 0.52 points. Activating prefixes of the 12-block stack, however, drives the worst ICL contribution to −16.09 points. The harm is therefore a property of composition across depth, not one obviously bad late block.

Starting from the shipped model and forcing block 0 to uniform improves every ICL configuration, materially in 8 of 10, by 0.4 to 20.5 points. Uniforming every block removes 11.8% of model parameters but improves ICL by 5.92 points on average; changing block 0 alone does better at 7.2.

06

What the experiment does—and does not—support

The clean lesson is operational: when example windows are not aligned, compare learned routing with a uniform-attention null and keep the univariate baseline. A learned attention map is not automatically more informative than a fixed pool. The broader claim “attention weights are useless” would be wrong: learned weighting helps 8 of 10 multivariate sensor-network configurations, and the learned V/O projection beats a plain cross-series mean in the paper's control.

Next, probe phase alignment directly, repeat the intervention on other architectures, and retrain the modified design rather than only patching a checkpoint. For the underlying mechanism, see the interactive Transformer walkthrough; the softmax and weighted-average pieces above are the same primitives, applied across series rather than tokens.

07

Reproduction notes

The charts labeled paper-reported transcribe Table A1 and the main-text counts. The temperature and three-value examples are deterministic, illustrative calculations from the post's pure core functions. No external assets, stochastic simulation, or ML library is used.

References

  1. Fore, M.; Inder, J. M.; Nair, M.; Vaddamanu, P.; Keshava, S. (2026). Pooling Helps, Learned Weighting Hurts In-Context: Decomposing Group Attention. arXiv:2610.01831
  2. Ansari, A. F. et al. (2025). Chronos-2: From Univariate to Universal Forecasting. arXiv:2510.15821