← Blog/blog/llm-booking-logit-scale-trap

The LLM preference coefficient that mostly measures decisiveness

01

The buyer is now a model

A booking agent can satisfy every explicit constraint and still make a consequential choice: the $148 room or the $281 room, a familiar chain or an independent hotel, the cheapest acceptable listing or the highest-rated one. PriceBench treats those choices as a revealed purchasing policy rather than merely asking whether the task completed.

Kireyev builds 3,600 tasks from 179 real New York City hotel profiles and scores 28 language models from eight providers. Each task is shown twice with its options swapped. That simple counterbalance matters: position bias can then enter its own intercept instead of masquerading as a taste for price or quality.

Paper-reported first-slot rates for the five locked models. The 85% bar is the preregistered engagement ceiling; the four distinct plotted rates all fail it.
02

Turn paired choices into a logit

For two hotels, the model predicts the probability of choosing A from three pieces: a first-slot intercept, the difference in log prices, and the difference in quality. A negative price coefficient means higher prices repel the chooser; a positive quality coefficient means better reviews attract it. The site runs the TypeScript below; Python and C++ are line-for-line translations.

def choice_diagnostics(price_a, price_b, quality_a, quality_b,
                       intercept, beta_log_price, beta_quality, reference_price):
    log_odds = (
        intercept
        + beta_log_price * math.log(price_a / price_b)
        + beta_quality * (quality_a - quality_b)
    )
    probability_a = 1 / (1 + math.exp(-log_odds))
    willingness_to_pay = (
        beta_quality / abs(beta_log_price) * reference_price
    )
    return probability_a, willingness_to_pay

Log price makes sensitivity proportional: a 10% price increase has the same utility effect near $100 and $800. The paper also fits linear price and price-decile specifications, because a seller needs to know whether a fixed-dollar discount, a proportional discount, or crossing a price threshold moves the agent.

soft chooser (scale 1×)decisive chooser (scale 3×)
Illustrative deterministic calculation from the core module. Both choosers trade one review point against price at the same break-even ratio; multiplying every coefficient by three only makes choices sharper.
03

The coefficient-scale trap

Here is the skim-reading mistake. A large price coefficient does not by itself mean “cares more about price.” In a logit, all coefficients share an unobserved scale. A model that responds more consistently can post large price and quality coefficients together even when the trade-off between them is unchanged. PriceBench finds exactly that diagonal: raw price and quality weights correlate at 0.55.

|price coefficient| / baselinequality coefficient / baselineWTP / baseline
Illustrative scale sweep from the core module. Both raw coefficients grow with the common logit scale; their ratio, expressed as willingness to pay (WTP), stays fixed.

With log price, the ratio has a concrete local interpretation at reference price p: WTP = βquality / |βlog price| × p. At the paper's $250 reference night, the 20 well-identified agents span about sixteen-fold—from $68 to $1,071 for one review-score point.

Modelβ log priceWTP / review pointWTP rank
Gemma2 9B−5.71$6820
GPT-5.4−7.02$10719
GPT-4.1 Mini−5.68$11918
GPT-5.4 Mini−4.56$22212
GPT-5.4 Nano−2.47$5005
Phi-4 Minipaper appendix$1,0711
Selected paper-reported appendix estimates. The Phi-4 Mini row is included for its reported WTP extreme; its coefficient is not transcribed here.
Selected paper-reported willingness-to-pay estimates at a $250 nightly reference price. These are fitted LLM-choice quantities, not human valuations or recommendations.
04

The ratio changes what gets bought

This is not just coefficient geometry. On the identical 3,600 binary tasks, GPT-5.4 books a mean $247 night while Phi-4 Mini books $393. The paper attributes the $146 gap to their different price–quality lean, after separating it from overall decisiveness.

ModelMean booked nightly pricePaper's reading
GPT-5.4$247strong price weight
Phi-4 Mini$393weak price, heavy quality weight
Paper-reported means on the same task set.
Paper-reported exact-hit rates when binary-estimated utilities predict a five-option choice. All beat the 20% chance bar; the test still reuses the same hotel pool.
05

What the benchmark does not identify

The headline estimates come from a neutral, forced-choice hotel prompt at temperature zero. Prompt variants, sampled decoding, and five-option lists preserve broad price–quality ordering on tested subsets, but the dollar levels move with format. Only six of 28 models enter the prompt ablation, human choices on these exact tasks were not collected, and tool-using multi-turn agents remain outside the evidence.

Brand coefficients need still more restraint. Controlling for observed hotel attributes cannot distinguish a true brand preference from a brand association absorbed during pretraining. Either can affect a booking, but they imply different interventions and explanations.

06

From the audit back to the model

The fitted choice equation is ordinary binary logistic regression on attribute differences. You can manipulate that machinery directly on the interactive logistic-regression page. The important extra lesson here is interpretive: before reading a coefficient as preference, ask what its scale is allowed to absorb.

References

  1. Pavel Kireyev (2026). PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents. EMNLP 2026 Industry Track / arXiv:2609.31468
  2. Daniel McFadden (1974). Conditional logit analysis of qualitative choice behavior. Frontiers in Econometrics
  3. Kenneth Train (2009). Discrete Choice Methods with Simulation. Cambridge University Press