Test-time compute sounds simple: run a tabular foundation model more times, then combine the answers. Ning and colleagues show why that slogan is incomplete. A pool of 96 configurations improves only after validation labels decide who gets a vote. Let everyone vote equally and error rises instead.
The ensemble is fitted after training
Each candidate is one prediction configuration: a different checkpoint, context construction, view, or adaptation choice. Greedy ensemble selection starts empty. For 40 rounds it tentatively adds every candidate, keeps the one with the lowest validation log loss, and allows repeated picks. Selection frequency becomes the final weight.
def select(labels, members, rounds):
counts = [0] * len(members)
total = [0.0] * len(labels)
for step in range(rounds):
losses = []
for member in members:
trial = [(s + p) / (step + 1) for s, p in zip(total, member)]
losses.append(log_loss(labels, trial))
best = min(range(len(losses)), key=lambda i: losses[i])
counts[best] += 1
total = [s + p for s, p in zip(total, members[best])]
return [count / rounds for count in counts]This is supervised fitting even though no model parameters move. The labels used by the selector therefore belong in the statistical budget, not merely in an engineering footnote.
| Aggregator | Relative error | 95% CI | Elo gain |
|---|---|---|---|
| Uniform average | +3.68% | [+0.02, +8.86] | — |
| Best single | −1.25% | [−3.41, +0.07] | +34 |
| Greedy selection | −2.44% | [−4.75, −0.80] | +67 |
One eighth of the labels cannot choose the crowd
The most important table is not the headline score. The authors rerun selection with fractions of the validation set. With only one eighth of the labels, every tested pool size remains worse than the default model. A larger pool shrinks the damage, but it does not cross zero. Only at one quarter does the 16-member pool become marginally useful.
| Labels | K=8 | K=16 | K=32 | K=64 | K=96 |
|---|---|---|---|---|---|
| 1/8 | +1.37 | +0.83 | +0.61 | +0.44 | +0.39 |
| 1/4 | +0.36 | −0.19 | −0.48 | −0.72 | −0.81 |
| 1/2 | −0.35 | −0.99 | −1.34 | −1.66 | −1.77 |
| all | −0.77 | −1.45 | −1.85 | −2.27 | −2.44 |
A toy selector makes the failure visible
The next chart is deterministic and illustrative, not a reproduction of a paper dataset. One candidate is reliably moderate. Another is nearly perfect on the first four validation rows and confidently wrong later. A third is cautious. The selector sees prefixes of 2, 4, 8, 12, and 16 labels; its weights are then scored on the same fixed 16-row evaluation set. The small-prefix validation score looks excellent precisely when the evaluation score is worst. The equally weighted toy ensemble has log loss 0.583.
The strongest gain is also the expensive one
Cheap native views deliver +21 Elo in 1.66 GPU-hours. The 96-configuration selection sweep delivers +67 Elo but consumes 247.5 GPU-hours; combining mechanisms reaches +93 Elo at 253.1 GPU-hours. A separate frozen TabFM baseline scores higher than that combined system while using 19.3 GPU-hours. Test-time scaling is therefore a mechanism-and-budget choice, not a generic promise that more inference wins.
Gains are uneven too: across methods, the five largest positive dataset gains account for 46–70% of all positive error reduction. The paper's attempt to predict which datasets benefit reaches cross-validated R² of only 0.012 at best. The average gain is real, but advance routing remains unreliable.
What to probe next
Reserve a second, untouched audit set for choosing the selector itself; plot gain against validation examples rather than fractions; compare greedy selection with shrinkage toward uniform or the default; and report compute-adjusted regret when candidate generation and label acquisition share one budget. Finally, test whether the chosen weights survive dataset shift.
References
- Kanghui Ning, Marin Biloš, James T. Wilson, Yilang Zhang, Kashif Rasul, Dongjin Song, Anderson Schneider, and Yuriy Nevmyvaka (2026). Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits. arXiv:2610.12005
- Rich Caruana, Alexandru Niculescu-Mizil, Geoff Crew, and Alex Ksikes (2004). Ensemble Selection from Libraries of Models. ICML 2004