deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (native)

Tabular foundation models beat tuned XGBoost on 14 datasets without training, benchmark finds

An independent RTX 4070 Ti benchmark found TabPFN and TabICL outscoring tuned XGBoost across all fourteen datasets tested, with no training on the target tables. The per-model record is more mixed on wide tables.

Tabular foundation models beat tuned XGBoost on 14 datasets without training, benchmark finds

An independent benchmark by Efraín Garay, surfaced on Hacker News's front page, put a months-old claim to a measurable test: that a tabular foundation model can predict on a table it has never trained on and still beat tuned gradient boosting. Run locally on an RTX 4070 Ti, the test covered fourteen datasets — and according to Garay, the no-training approach finished ahead of hyperparameter-tuned XGBoost on all of them.

How the test was built

Every dataset came from the inria-soda tabular benchmark suite, trimmed to 3,000 rows and split 70/30 with stratification at a fixed seed. Four contenders shared the same stopwatch; the write-up details TabICL, TabPFN and tuned XGBoost, with untuned XGBoost appearing as a floor. Eight of the fourteen datasets are itemised, spanning credit risk, sales campaigns, hospital readmission, recidivism scoring, electricity demand and molecular screening.

What a tabular foundation model actually does

TabPFN and TabICL are pretrained on millions of purpose-built synthetic tables. When a new table arrives, no weight is updated: the training rows — 2,100 in these tests — are passed in as context, and predictions come back in a single forward pass. It is the in-context learning mechanism known from language models, moved from tokens to rows and columns.

Garay notes that the familiar fit call still exists in the code, but internally it does nothing more than copy data to the GPU. That flips the cost profile engineers expect: fitting is nearly free and prediction is where compute concentrates — the mirror image of gradient-boosted trees, where training dominates.

The results, rough edges included

On the eight itemised datasets, TabICL won seven and tied the eighth. In credit risk, TabPFN scored 0.7300 on heloc against 0.7078 for tuned XGBoost, and TabICL took the credit table 0.7667 to 0.7533. But neither model swept the field. On electricity demand, TabICL matched tuned XGBoost to the fourth decimal (0.8178) while TabPFN fell below at 0.8111 — a tie, not a win. On Bioresponse, a 419-column table of chemical descriptors, the two foundation models split: TabICL led at 0.7922 versus 0.7744, but TabPFN dropped to 0.7567, below even untuned XGBoost, and needed 18.1 seconds to do it. Garay's conclusion: table width, not length, is what squeezes these models.

Accuracy also hides part of the story. On the Diabetes130US readmission data, accuracies were nearly identical, but the AUC — the ranking quality a triage workflow actually relies on — separated clearly: 0.6484 for TabICL against 0.6264 for tuned XGBoost. TabICL's single-pass prediction times ran 0.6 to 0.8 seconds on the narrow tables and 6.0 seconds on Bioresponse.

Garay attaches his own caveat to compas-two-years, the dataset behind the debate over algorithmic bias in the courts: a better score does not make the recidivism use case legitimate, because that argument was never about accuracy.

Why it matters

Standard tabular practice assumes a tuning loop: search hyperparameters, fit, evaluate, repeat. If a pretrained model can match or beat that loop on small and medium tables simply by being handed the rows, hyperparameter search stops being a mandatory step and becomes an occasional refinement, and engineering attention shifts from optimiser settings to data quality and evaluation design.

The limits matter as much as the headline. This is a single-author benchmark on one consumer GPU, capped at 3,000 rows per table, and the per-model record frays on wide tables — TabPFN's especially. The defensible takeaway is not that XGBoost is obsolete, but that on tables of this size it now has a serious zero-training rival that deserves a place in the comparison before any tuning begins.

  • #machine-learning
  • #tabular-data
  • #xgboost
  • #foundation-models
  • #benchmarks

Related posts