· via dev.to (home feed)
Training-free TabPFN and TabICL beat tuned XGBoost on all 14 benchmark tables
A dev.to benchmark pits the in-context tabular models TabPFN and TabICL against tuned XGBoost on 14 Grinsztajn datasets, and the models that skip training won every table.

What the benchmark found
A practitioner benchmark by Efrain Garay, published on dev.to, pitted two in-context tabular models, TabPFN and TabICL, against a tuned XGBoost setup on 14 datasets taken from the Grinsztajn benchmark, a widely used suite of small- to medium-sized tabular classification and regression tasks. According to Garay, the models that never trained on those tables won on 14 out of 14.
The scoreline matters less than the mechanism behind it. TabPFN and TabICL belong to a family of tabular foundation models that skip the fitting step entirely: they are pre-trained elsewhere and produce predictions for a new dataset by reading it as input, not by running gradient descent on it. XGBoost, by contrast, was tuned for the comparison and still lost every table.
How the comparison was set up
Garay writes that every method was evaluated on the same data splits and under the same measurement conditions. The 14 datasets come from the Grinsztajn suite, the collection introduced by a 2022 study of why tree-based models still beat deep learning on typical tabular data. That makes the venue pointed: the benchmark built to demonstrate boosting's dominance on ordinary tables is where the training-free models swept it. The full measurements, including per-dataset results and resource costs, sit in a longer write-up on the author's blog.
What training-free prediction means here
TabPFN is a prior-fitted network: a transformer pre-trained on a large corpus of synthetic datasets. Handed a new table, it treats the training rows as context and emits predictions directly, with no epochs, no per-dataset gradients and little of the hyperparameter search that dominates applied tabular work. TabPFN v2 was described in a Nature paper in early 2025, which reported it beating tuned baselines on small datasets. TabICL, introduced in 2025 by a Google-led research team, extends the same idea with a larger and more varied pre-training corpus, aiming at bigger tables and at regression as well as classification.
The practical consequence is a change of workflow rather than just accuracy: a strong tabular predictor becomes something you load from a checkpoint, closer to calling a model than to owning a training pipeline.
Where boosting keeps its edge
Garay's post also reports what the training-free approach costs in latency and VRAM, the resources in-context prediction spends instead of a training budget, and describes the situations where he would still reach for boosting. Because these models read the training table as context at prediction time, their compute and memory usage scale with the table rather than being paid off once during training. That points to the natural fault line: small and medium tables, where the sweep happened, favour the foundation models, while very large tables and tight-latency serving remain boosting territory. The exact latency and memory figures are in the author's full write-up rather than the dev.to summary.
Why it matters
Gradient-boosted trees have been the default answer for tabular data for roughly a decade, and tuning an XGBoost is the reflex first move in most applied machine learning. If a pre-trained model can be dropped onto an unseen table with no tuning loop and beat that reflex on every dataset in a standard benchmark suite, the sensible baseline changes: try the foundation model first, keep boosting for the cases it handles poorly. One practitioner-run comparison on 14 datasets does not settle the question, and both the resource costs and the single-author methodology warrant scrutiny, but the result points the same way as the peer-reviewed literature on these models: tabular foundation models are moving from curiosity toward credible default.
- #tabular-data
- #machine-learning
- #xgboost
- #tabpfn
- #benchmark