← Blog/blog/tabular-ensemble-validation-budget

The 96-model ensemble that gets worse when everyone votes

Test-time compute sounds simple: run a tabular foundation model more times, then combine the answers. Ning and colleagues show why that slogan is incomplete. A pool of 96 configurations improves only after validation labels decide who gets a vote. Let everyone vote equally and error rises instead.

01

The ensemble is fitted after training

Each candidate is one prediction configuration: a different checkpoint, context construction, view, or adaptation choice. Greedy ensemble selection starts empty. For 40 rounds it tentatively adds every candidate, keeps the one with the lowest validation log loss, and allows repeated picks. Selection frequency becomes the final weight.

def select(labels, members, rounds):
    counts = [0] * len(members)
    total = [0.0] * len(labels)
    for step in range(rounds):
        losses = []
        for member in members:
            trial = [(s + p) / (step + 1) for s, p in zip(total, member)]
            losses.append(log_loss(labels, trial))
        best = min(range(len(losses)), key=lambda i: losses[i])
        counts[best] += 1
        total = [s + p for s, p in zip(total, members[best])]
    return [count / rounds for count in counts]

This is supervised fitting even though no model parameters move. The labels used by the selector therefore belong in the statistical budget, not merely in an engineering footnote.

Paper-reported relative error change for the 96-configuration pool; lower is better. The negative bars are improvements. Values and confidence intervals are from the paper, not recomputed here.
AggregatorRelative error95% CIElo gain
Uniform average+3.68%[+0.02, +8.86]—
Best single−1.25%[−3.41, +0.07]+34
Greedy selection−2.44%[−4.75, −0.80]+67
Paper-reported aggregation results for K = 96. Best-single and greedy reuse validation labels; uniform averaging does not.
02

One eighth of the labels cannot choose the crowd

The most important table is not the headline score. The authors rerun selection with fractions of the validation set. With only one eighth of the labels, every tested pool size remains worse than the default model. A larger pool shrinks the damage, but it does not cross zero. Only at one quarter does the 16-member pool become marginally useful.

K = 8K = 32K = 96
Paper Table 14. Relative error change versus the default across selector-label fractions 1/8, 1/4, 1/2, and all. Lower is better; zero is the break-even line.
LabelsK=8K=16K=32K=64K=96
1/8+1.37+0.83+0.61+0.44+0.39
1/4+0.36−0.19−0.48−0.72−0.81
1/2−0.35−0.99−1.34−1.66−1.77
all−0.77−1.45−1.85−2.27−2.44
Paper-reported relative error changes (%). Positive means worse than the default configuration.
03

A toy selector makes the failure visible

The next chart is deterministic and illustrative, not a reproduction of a paper dataset. One candidate is reliably moderate. Another is nearly perfect on the first four validation rows and confidently wrong later. A third is cautious. The selector sees prefixes of 2, 4, 8, 12, and 16 labels; its weights are then scored on the same fixed 16-row evaluation set. The small-prefix validation score looks excellent precisely when the evaluation score is worst. The equally weighted toy ensemble has log loss 0.583.

selector validation lossfixed evaluation loss
Seed-free illustrative selection. Prefix-specialist predictions win tiny validation prefixes; representative labels shift weight toward the robust candidate. Both curves are computed by the pure TypeScript core used in the code walkthrough.
04

The strongest gain is also the expensive one

Cheap native views deliver +21 Elo in 1.66 GPU-hours. The 96-configuration selection sweep delivers +67 Elo but consumes 247.5 GPU-hours; combining mechanisms reaches +93 Elo at 253.1 GPU-hours. A separate frozen TabFM baseline scores higher than that combined system while using 19.3 GPU-hours. Test-time scaling is therefore a mechanism-and-budget choice, not a generic promise that more inference wins.

Paper-reported total GPU-hours over 51 datasets. Bars use a linear scale; the aggregation search dominates the compute bill.

Gains are uneven too: across methods, the five largest positive dataset gains account for 46–70% of all positive error reduction. The paper's attempt to predict which datasets benefit reaches cross-validated R² of only 0.012 at best. The average gain is real, but advance routing remains unreliable.

05

What to probe next

Reserve a second, untouched audit set for choosing the selector itself; plot gain against validation examples rather than fractions; compare greedy selection with shrinkage toward uniform or the default; and report compute-adjusted regret when candidate generation and label acquisition share one budget. Finally, test whether the chosen weights survive dataset shift.

References

  1. Kanghui Ning, Marin Biloš, James T. Wilson, Yilang Zhang, Kashif Rasul, Dongjin Song, Anderson Schneider, and Yuriy Nevmyvaka (2026). Test-Time Compute for Tabular Foundation Models: Mechanisms, Gains, and Limits. arXiv:2610.12005
  2. Rich Caruana, Alexandru Niculescu-Mizil, Geoff Crew, and Alex Ksikes (2004). Ensemble Selection from Libraries of Models. ICML 2004