A market turns a price ladder into a forecast
Kalshi lists mutually exclusive contracts for tomorrow's high temperature: perhaps 81°F or lower, 82–83°F, and so on. A contract paying one dollar in its interval has a price that can be read as a probability. Crosier joins those rungs into a distribution, then compares its implied mean with six public weather products over 7,590 city-days in seven U.S. cities.
The paper's headline is striking. At the first hour of trading, the market has about 10% lower root-mean-square error, or RMSE, than the National Blend of Models (NBM), the strongest single public product. RMSE squares each forecast error, averages those squares, then takes the square root, so large misses receive extra weight.
The published single-model result is real
| Anchor | Market RMSE | NBM RMSE | Reported edge |
|---|---|---|---|
| Opening hour, d−1 | 2.44°F | 2.70°F | 9.8% |
| Evening, d−1 | 2.21°F | 2.45°F | 10.0% |
| Final bulletin | 1.88°F | 2.12°F | 11.4% |
Prices are not automatically a clean probability distribution. The paper uses bid–ask midpoints and screens ladders whose quoted mass is too sparse, too wide, or more than 30% away from one. Our pure function normalizes the mass before taking Σpᵢcᵢ. For the illustrative quotes [0.20, 0.50, 0.40] on centers [70, 72, 74], the normalized mean is 72.36°F.
A stronger ensemble absorbs most of the headline
The paper also constructs a forecast from every public product. Each city is fit only on earlier days, so tomorrow never leaks into today's weights. This rolling linear regression is a “free rider”: it sees the same public stack a sophisticated trader could read. The market still wins, but the margin is much narrower.
def relative_rmse_advantage(market_rmse, benchmark_rmse):
if market_rmse < 0 or benchmark_rmse <= 0:
raise ValueError('invalid RMSE')
return 100 * (benchmark_rmse - market_rmse) / benchmark_rmseLiquidity screens matter most at the fragile edge
Three quote-quality screens remove 14% of observations. Without them, the market still beats NBM by 8–9% and every anchor remains significant at 1%. Against the public combination, though, the final-bulletin edge falls from 3.0% to 0.7% and is no longer statistically significant. That is the place where a clean headline becomes a conditional finding.
The authors run sensible robustness checks. Using the implied median avoids assumptions about open temperature tails, and headline gaps move by less than two percentage points. They also cluster the comparison by date and use Newey–West errors because nearby cities and consecutive weather systems do not create independent observations.
The public forecast moves farther toward the price
The event study asks which signal leads. A typical market–NBM disagreement is 1.49°F and the pooled revision slope is 0.042, implying a 0.063°F NBM move. A typical NBM revision is 0.43°F and the market-response slope is 0.037, implying a 0.016°F price move. Those are regression associations, not proof that forecasters literally copy traders.
What to probe next
Evaluate a predeclared equal-weight public ensemble on exactly the same days as every market quote; report calibration and full-distribution scores such as log loss alongside the mean's RMSE; stratify by traded volume and spread; and test whether the residual edge survives transaction costs for anyone who must trade to read it. A live prospective holdout would also separate a durable information advantage from one historical market regime.
For the combination mechanism, continue with the linear-regression walkthrough. For another example of a seemingly decisive forecast result changing under data provenance, compare the pretrained forecasting look-ahead audit.
References
- A. W. Crosier (2026). Prediction Markets Beat the Weather Forecast on Tomorrow's High Temperature. arXiv:2609.23969 [q-fin.GN]
- F. X. Diebold and R. S. Mariano (1995). Comparing Predictive Accuracy. Journal of Business & Economic Statistics 13(3)