The buyer is now a model
A booking agent can satisfy every explicit constraint and still make a consequential choice: the $148 room or the $281 room, a familiar chain or an independent hotel, the cheapest acceptable listing or the highest-rated one. PriceBench treats those choices as a revealed purchasing policy rather than merely asking whether the task completed.
Kireyev builds 3,600 tasks from 179 real New York City hotel profiles and scores 28 language models from eight providers. Each task is shown twice with its options swapped. That simple counterbalance matters: position bias can then enter its own intercept instead of masquerading as a taste for price or quality.
Turn paired choices into a logit
For two hotels, the model predicts the probability of choosing A from three pieces: a first-slot intercept, the difference in log prices, and the difference in quality. A negative price coefficient means higher prices repel the chooser; a positive quality coefficient means better reviews attract it. The site runs the TypeScript below; Python and C++ are line-for-line translations.
def choice_diagnostics(price_a, price_b, quality_a, quality_b,
intercept, beta_log_price, beta_quality, reference_price):
log_odds = (
intercept
+ beta_log_price * math.log(price_a / price_b)
+ beta_quality * (quality_a - quality_b)
)
probability_a = 1 / (1 + math.exp(-log_odds))
willingness_to_pay = (
beta_quality / abs(beta_log_price) * reference_price
)
return probability_a, willingness_to_payLog price makes sensitivity proportional: a 10% price increase has the same utility effect near $100 and $800. The paper also fits linear price and price-decile specifications, because a seller needs to know whether a fixed-dollar discount, a proportional discount, or crossing a price threshold moves the agent.
The coefficient-scale trap
Here is the skim-reading mistake. A large price coefficient does not by itself mean “cares more about price.” In a logit, all coefficients share an unobserved scale. A model that responds more consistently can post large price and quality coefficients together even when the trade-off between them is unchanged. PriceBench finds exactly that diagonal: raw price and quality weights correlate at 0.55.
With log price, the ratio has a concrete local interpretation at reference price p: WTP = βquality / |βlog price| × p. At the paper's $250 reference night, the 20 well-identified agents span about sixteen-fold—from $68 to $1,071 for one review-score point.
| Model | β log price | WTP / review point | WTP rank |
|---|---|---|---|
| Gemma2 9B | −5.71 | $68 | 20 |
| GPT-5.4 | −7.02 | $107 | 19 |
| GPT-4.1 Mini | −5.68 | $119 | 18 |
| GPT-5.4 Mini | −4.56 | $222 | 12 |
| GPT-5.4 Nano | −2.47 | $500 | 5 |
| Phi-4 Mini | paper appendix | $1,071 | 1 |
The ratio changes what gets bought
This is not just coefficient geometry. On the identical 3,600 binary tasks, GPT-5.4 books a mean $247 night while Phi-4 Mini books $393. The paper attributes the $146 gap to their different price–quality lean, after separating it from overall decisiveness.
| Model | Mean booked nightly price | Paper's reading |
|---|---|---|
| GPT-5.4 | $247 | strong price weight |
| Phi-4 Mini | $393 | weak price, heavy quality weight |
What the benchmark does not identify
The headline estimates come from a neutral, forced-choice hotel prompt at temperature zero. Prompt variants, sampled decoding, and five-option lists preserve broad price–quality ordering on tested subsets, but the dollar levels move with format. Only six of 28 models enter the prompt ablation, human choices on these exact tasks were not collected, and tool-using multi-turn agents remain outside the evidence.
Brand coefficients need still more restraint. Controlling for observed hotel attributes cannot distinguish a true brand preference from a brand association absorbed during pretraining. Either can affect a booking, but they imply different interventions and explanations.
From the audit back to the model
The fitted choice equation is ordinary binary logistic regression on attribute differences. You can manipulate that machinery directly on the interactive logistic-regression page. The important extra lesson here is interpretive: before reading a coefficient as preference, ask what its scale is allowed to absorb.
References
- Pavel Kireyev (2026). PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents. EMNLP 2026 Industry Track / arXiv:2609.31468
- Daniel McFadden (1974). Conditional logit analysis of qualitative choice behavior. Frontiers in Econometrics
- Kenneth Train (2009). Discrete Choice Methods with Simulation. Cambridge University Press