← Blog/blog/agent-token-reread-tax

The token forecast that improves after the bill starts

01

An agent pays for old words again

A chat completion bill is not just the new text an agent writes. Each later model call also sends the context retained from earlier calls. A tool result that adds 600 tokens can therefore cost far more than 600: if five later calls reread it, the downstream bill grows by 3,000 input tokens before counting any new output.

TokenCast studies that moving target across 11,712 traces from four task suites, two agent harnesses, and six language models. The same task can take a different path on each run, so a useful forecast must update as tool failures, retries, and generated prefixes arrive.

02

The cross-term is the whole trick

The paper represents one contiguous execution segment by three numbers: call count n, net context growth g, and residual cost b. Starting from input length L, its cost is nL + b. When segment B follows A, B inherits A's added context, so composition needs the extra term nBgA.

def compose(first, second):
    n_a, g_a, b_a = first
    n_b, g_b, b_b = second
    reread_tax = n_b * g_a
    return (
        n_a + n_b,
        g_a + g_b,
        b_a + b_b + reread_tax,
    )

The site runs the TypeScript above; Python and C++ are faithful translations. In the illustrative example below, the ordinary sum stays flat because it prices both segments from the original 1,000-token boundary. Exact composition rises with the context that later calls must reread. At 600 tokens of prefix growth, the omitted cross-term is 3,000 tokens.

exact composed billnaive segment sum
Illustrative deterministic calculation from the core module. Segment B has five calls; every token added by A is therefore billed five more times. These are synthetic inputs, not paper measurements.
03

A small boundary miss can become a large bill miss

The same identity explains the paper's weighted training loss. If a predictor misses prefix growth by 120 tokens, the downstream cost error is 120 times the number of suffix calls. That is why TokenCast weights a growth error by the future call count instead of treating every local error equally.

cost error from a 120-token growth miss
Illustrative error propagation computed by the core module. The growth error is fixed at 120 tokens; only the number of later calls changes.
04

The 14.5% headline starts too early

Across 96 benchmark–model–prediction-point comparisons, TokenCast lowers mean absolute error against the strongest comparator by 14.5% on average. But the gain is not evenly available. At task start—before the agent has done anything—the method is 2.2% worse on average and loses 15 of 24 comparisons. The large gains arrive after execution evidence exists: 30.4% during generation and 27.8% after completed calls.

Paper-reported average MAE relative to the strongest comparator at each prediction point. Lower is better; 100 means a tie. Task start is 102.2 because TokenCast is 2.2% worse on average.
Prediction pointTokenCast MAEStrongest comparator MAEReduction
Task start144.0k152.0k5.3%
Call start64.569.57.2%
In-call update38.974.647.9%
Task update80.0k115.0k30.4%
Paper-reported SWE-bench Verified results for GPT-5.4. Task-level errors use thousands of tokens; call-level errors use tokens, so compare within rows rather than across rows.
05

Intervals reveal the remaining uncertainty

The authors calibrate nominal 90% intervals on held-out tasks. At task start, TokenCast covers 82.0% of repeated outcomes—not 90%—while the uncalibrated Self-Prediction intervals cover 52.7%. TokenCast's intervals are also more than twice as wide, a reminder that improved coverage is partly purchased with honest uncertainty.

Paper-reported task-start interval coverage on repeated anchor-task runs. The methods are not calibrated identically: TokenCast uses held-out calibration; Self-Prediction contributes native intervals.

In offline budget replay, the forecast-driven policy consumes 21.3% fewer tokens than fixed budgets at matched trace completion. That last phrase matters: replay checks whether a recorded run reaches its terminal state, not whether the underlying software task is solved.

Paper-reported average token saving across seven replay budgets, indexed to the fixed-budget policy. Trace completion is matched; task success is not measured by this replay.
06

What to probe next

Zero-shot domain transfer is still fragile: on independently released LiveClawBench trajectories, error ratios are 1.31 at call start and 1.47 at task update relative to Self-Prediction; twenty target-domain tasks bring both below one. A useful next audit would separate improvements from the segment identity itself, the hand-engineered execution features, and target-domain adaptation—and test real task success under online stopping rather than replay alone.

TokenCast fits its direct, prefix, suffix, and correction predictors with gradient-boosted trees. You can inspect the histogram and leaf-wise mechanics on the interactive LightGBM page. The transferable idea is simpler than the predictor: forecast the state passed across a boundary, because downstream cost depends on it.

References

  1. Chaoqian Ouyang, Ling Yue, Libin Zheng, Huanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, and Shimin Di (2026). TokenCast: Forecasting Token Consumption During LLM Agent Execution. arXiv:2609.35760
  2. Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, and Ofir Press (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. ICLR 2024
  3. Tilmann Gneiting and Adrian E. Raftery (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association