← Blog/blog/test-time-scaling-schedule-budget

Eight LLM samples can cost five times more

01

Eight answers are not one systems budget

Test-time scaling asks a model for several candidate answers and then chooses among them. A common version is self-consistency: sample independent reasoning paths and return the plurality answer. The candidate count N sounds like a complete budget. It is not.

Eight candidates can arrive in one call as a batch of eight, two calls of four, four calls of two, or eight serial calls of one. Kashaniyan and Jannesari hold N at eight and change only that schedule. On A100 GPUs, the serial endpoint consumes 4.64× the gross device energy for Phi-3 and 4.86× for Qwen, while P95 latency grows 5.77× and 6.12×.

02

Write the schedule, not just N

Let S = (b₁,…,bC), where bc is the number of candidates in call c. Then N = Σbc, while C records how many times the inference engine is entered. All four schedules below generate eight logical candidates and repeat a 125-token prompt 1,000 logical token-times. Their execution shapes are different.

ScheduleCallsCandidates / callLogical input tokens
1 × 8181,000
2 × 4241,000
4 × 2421,000
8 × 1811,000
Deterministic accounting example with a 125-token prompt; no measured GPU values are used.
Phi-3 energyQwen energyPhi-3 P95 latencyQwen P95 latency
Paper-reported fixed-N=8 systems costs, normalized within model to one call of eight. X positions are 1×8, 2×4, 4×2, and 8×1.

Why does serial execution lose? A larger batch exposes more independent decoding work to the GPU at once. Smaller calls repeat call overhead and prompt processing while offering less concurrent work. The experiment measures the end-to-end result; it does not separately identify how much each mechanism contributes.

03

Per-token efficiency can improve while the query gets dearer

The paper also increases batched N from one to eight. Energy per generated token falls because the batch uses the device more efficiently, yet total joules per query rise because the query asks for more output. Those statements are compatible: a cheaper token can still produce a more expensive request.

Paper-reported gross GPU-device joules per query for batched N. Phi-3 ran on V100 and Qwen on A100, so compare each model with itself, not against the other model.
Phi-3 / V100Qwen / A100
Paper-reported joules per generated token at N = 1, 2, 4, 8. Efficiency improves even as total query energy rises.
NPhi-3 J/queryQwen J/queryPhi-3 J/tokenQwen J/token
16315962.6342.222
27956701.6071.265
49748021.0040.758
812869750.6550.459
Table V of the paper; means over three repetitions. Hardware differs by model in this experiment.
04

Two samples can vote exactly like one

The accuracy curve contains a smaller but revealing protocol trap. Ties are resolved by the earliest candidate. With two disagreeing candidates, both have one vote, so the first answer wins. Therefore N=2 is mechanically identical to N=1 for every prompt—not merely similar on average.

Paper-reported tie frequency at N=2 on 500 GSM8K prompts. Earliest-wins tie handling makes primary N=2 accuracy equal N=1: 81.4% for Phi-3 and 51.4% for Qwen.

In the deterministic toy below, the first two candidates disagree on both prompts. Accuracy stays at 0% from N=1 to N=2, then jumps to 100% at N=3 when a majority forms. The first toy prompt returns A → A → B as N grows from one to three.

toy prefix accuracy
Illustrative deterministic prefixes, computed by the same earliest-wins plurality function used in the walkthrough; not paper data.
05

The energy ratio has a boundary

Energy is the difference between synchronized NVML cumulative device-energy readings around the measured query interval. It includes generation, answer extraction, and voting, but excludes model loading, warm-up, reporting-time token counting, and final grading. Idle GPU power is not subtracted, and CPU, memory, cooling, and the rest of the node are outside the boundary.

The fixed-N schedules also sample candidates independently. Their mean generated-token volume differs by only 0.8% for Phi-3 and 1.0% for Qwen, and length-stratified analyses preserve the pattern, but this is not a token-identical replay. The authors correctly frame the result as the cost of the complete schedule.

def summarize(schedule, energy_j, latency_s, generated_tokens):
    if not schedule or any(b <= 0 for b in schedule):
        raise ValueError('positive batch sizes required')
    candidates = sum(schedule)
    return {
        'candidates': candidates,
        'calls': len(schedule),
        'max_batch': max(schedule),
        'average_power_w': energy_j / latency_s,
        'candidates_per_s': candidates / latency_s,
        'joules_per_token': energy_j / generated_tokens,
    }

def plurality_vote(answers):
    counts = {answer: answers.count(answer) for answer in answers}
    return max(answers, key=lambda answer: counts[answer])
06

What the result does—and does not—authorize

The practical rule is conditional: for independent candidates, when memory permits, use fewer calls with larger batches. It need not survive unchanged under continuous batching, dependent search, speculative decoding, multi-GPU inference, larger models, or strict serving constraints. The exact ratios come from two small models, A100/V100 GPUs, and Hugging Face batch-scheduled generation.

A useful next experiment would replay identical sampled continuations or controlled token lengths across schedules, then decompose prompt processing, launch overhead, decoding utilization, and idle draw. A serving study should also add concurrent traffic and memory limits—the conditions under which one giant batch may be impossible or harmful to queueing latency.

For the model-side mechanics behind batched attention and decoding, continue with the Transformer walkthrough. The systems lesson is simpler: candidate count describes the logical search; schedule describes how the machine pays for it.