An agent pays for old words again
A chat completion bill is not just the new text an agent writes. Each later model call also sends the context retained from earlier calls. A tool result that adds 600 tokens can therefore cost far more than 600: if five later calls reread it, the downstream bill grows by 3,000 input tokens before counting any new output.
TokenCast studies that moving target across 11,712 traces from four task suites, two agent harnesses, and six language models. The same task can take a different path on each run, so a useful forecast must update as tool failures, retries, and generated prefixes arrive.
The cross-term is the whole trick
The paper represents one contiguous execution segment by three numbers: call count n, net context growth g, and residual cost b. Starting from input length L, its cost is nL + b. When segment B follows A, B inherits A's added context, so composition needs the extra term nBgA.
def compose(first, second):
n_a, g_a, b_a = first
n_b, g_b, b_b = second
reread_tax = n_b * g_a
return (
n_a + n_b,
g_a + g_b,
b_a + b_b + reread_tax,
)The site runs the TypeScript above; Python and C++ are faithful translations. In the illustrative example below, the ordinary sum stays flat because it prices both segments from the original 1,000-token boundary. Exact composition rises with the context that later calls must reread. At 600 tokens of prefix growth, the omitted cross-term is 3,000 tokens.
A small boundary miss can become a large bill miss
The same identity explains the paper's weighted training loss. If a predictor misses prefix growth by 120 tokens, the downstream cost error is 120 times the number of suffix calls. That is why TokenCast weights a growth error by the future call count instead of treating every local error equally.
The 14.5% headline starts too early
Across 96 benchmark–model–prediction-point comparisons, TokenCast lowers mean absolute error against the strongest comparator by 14.5% on average. But the gain is not evenly available. At task start—before the agent has done anything—the method is 2.2% worse on average and loses 15 of 24 comparisons. The large gains arrive after execution evidence exists: 30.4% during generation and 27.8% after completed calls.
| Prediction point | TokenCast MAE | Strongest comparator MAE | Reduction |
|---|---|---|---|
| Task start | 144.0k | 152.0k | 5.3% |
| Call start | 64.5 | 69.5 | 7.2% |
| In-call update | 38.9 | 74.6 | 47.9% |
| Task update | 80.0k | 115.0k | 30.4% |
Intervals reveal the remaining uncertainty
The authors calibrate nominal 90% intervals on held-out tasks. At task start, TokenCast covers 82.0% of repeated outcomes—not 90%—while the uncalibrated Self-Prediction intervals cover 52.7%. TokenCast's intervals are also more than twice as wide, a reminder that improved coverage is partly purchased with honest uncertainty.
In offline budget replay, the forecast-driven policy consumes 21.3% fewer tokens than fixed budgets at matched trace completion. That last phrase matters: replay checks whether a recorded run reaches its terminal state, not whether the underlying software task is solved.
What to probe next
Zero-shot domain transfer is still fragile: on independently released LiveClawBench trajectories, error ratios are 1.31 at call start and 1.47 at task update relative to Self-Prediction; twenty target-domain tasks bring both below one. A useful next audit would separate improvements from the segment identity itself, the hand-engineered execution features, and target-domain adaptation—and test real task success under online stopping rather than replay alone.
TokenCast fits its direct, prefix, suffix, and correction predictors with gradient-boosted trees. You can inspect the histogram and leaf-wise mechanics on the interactive LightGBM page. The transferable idea is simpler than the predictor: forecast the state passed across a boundary, because downstream cost depends on it.
References
- Chaoqian Ouyang, Ling Yue, Libin Zheng, Huanghui Guo, Shengxiang Xu, YiShu Wang, Ran Li, Jian Yin, Shaowu Pan, and Shimin Di (2026). TokenCast: Forecasting Token Consumption During LLM Agent Execution. arXiv:2609.35760
- Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, and Ofir Press (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. ICLR 2024
- Tilmann Gneiting and Adrian E. Raftery (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association