← Blog/blog/prediction-market-weather-benchmark

The weather market wins—until the benchmark gets smarter

01

A market turns a price ladder into a forecast

Kalshi lists mutually exclusive contracts for tomorrow's high temperature: perhaps 81°F or lower, 82–83°F, and so on. A contract paying one dollar in its interval has a price that can be read as a probability. Crosier joins those rungs into a distribution, then compares its implied mean with six public weather products over 7,590 city-days in seven U.S. cities.

The paper's headline is striking. At the first hour of trading, the market has about 10% lower root-mean-square error, or RMSE, than the National Blend of Models (NBM), the strongest single public product. RMSE squares each forecast error, averages those squares, then takes the square root, so large misses receive extra weight.

02

The published single-model result is real

AnchorMarket RMSENBM RMSEReported edge
Opening hour, d−12.44°F2.70°F9.8%
Evening, d−12.21°F2.45°F10.0%
Final bulletin1.88°F2.12°F11.4%
Paper-reported pooled results. d−1 means the day before the temperature is observed.
Kalshi marketNBM
Fig 1. Paper-reported RMSE at the opening hour, prior evening, and final NBM bulletin. Lower is better. The core implementation recomputes the unrounded gaps as 9.6%, 9.8%, 11.3%.

Prices are not automatically a clean probability distribution. The paper uses bid–ask midpoints and screens ladders whose quoted mass is too sparse, too wide, or more than 30% away from one. Our pure function normalizes the mass before taking Σpᵢcᵢ. For the illustrative quotes [0.20, 0.50, 0.40] on centers [70, 72, 74], the normalized mean is 72.36°F.

03

A stronger ensemble absorbs most of the headline

The paper also constructs a forecast from every public product. Each city is fit only on earlier days, so tomorrow never leaks into today's weights. This rolling linear regression is a “free rider”: it sees the same public stack a sophisticated trader could read. The market still wins, but the margin is much narrower.

def relative_rmse_advantage(market_rmse, benchmark_rmse):
    if market_rmse < 0 or benchmark_rmse <= 0:
        raise ValueError('invalid RMSE')
    return 100 * (benchmark_rmse - market_rmse) / benchmark_rmse
Fig 2. Opening-hour market advantage under three benchmarks. The 9.8% and 3.1% bars are paper-reported. The 2.1% bar is computed from the paper's equal-weight RMSE of 2.43°F and matched-sample market RMSE of 2.38°F. It is derived from reported values, not a new experiment.
04

Liquidity screens matter most at the fragile edge

Three quote-quality screens remove 14% of observations. Without them, the market still beats NBM by 8–9% and every anchor remains significant at 1%. Against the public combination, though, the final-bulletin edge falls from 3.0% to 0.7% and is no longer statistically significant. That is the place where a clean headline becomes a conditional finding.

NBM, screenedNBM, all quotesCombination, screenedCombination, all quotes
Fig 3. Paper-reported percentage advantages with and without liquidity screens at the same three anchors. The chart displays reported results; no missing market observations are imputed.

The authors run sensible robustness checks. Using the implied median avoids assumptions about open temperature tails, and headline gaps move by less than two percentage points. They also cluster the comparison by date and use Newey–West errors because nearby cities and consecutive weather systems do not create independent observations.

05

The public forecast moves farther toward the price

The event study asks which signal leads. A typical market–NBM disagreement is 1.49°F and the pooled revision slope is 0.042, implying a 0.063°F NBM move. A typical NBM revision is 0.43°F and the market-response slope is 0.037, implying a 0.016°F price move. Those are regression associations, not proof that forecasters literally copy traders.

Fig 4. Typical movement from the paper's pooled slopes and predictor standard deviations. NBM moves 3.9× as far toward the market as the market moves toward NBM. Both values are deterministic products of paper-reported inputs.
06

What to probe next

Evaluate a predeclared equal-weight public ensemble on exactly the same days as every market quote; report calibration and full-distribution scores such as log loss alongside the mean's RMSE; stratify by traded volume and spread; and test whether the residual edge survives transaction costs for anyone who must trade to read it. A live prospective holdout would also separate a durable information advantage from one historical market regime.

For the combination mechanism, continue with the linear-regression walkthrough. For another example of a seemingly decisive forecast result changing under data provenance, compare the pretrained forecasting look-ahead audit.

References

  1. A. W. Crosier (2026). Prediction Markets Beat the Weather Forecast on Tomorrow's High Temperature. arXiv:2609.23969 [q-fin.GN]
  2. F. X. Diebold and R. S. Mariano (1995). Comparing Predictive Accuracy. Journal of Business & Economic Statistics 13(3)