PFNs use native numeric architectures, while post-trained LLMs lag on larger tables

GaelVaroquaux · x · 2026-07-26

Gael Varoquaux explains that PFNs do not post-train LLMs. Instead, they use architectures that natively handle numbers without tokenizing them, and they train from scratch.

He contrasts this with papers like TabuLa-8B, saying that post-trained LLMs are only toy-scale solutions that work on very small datasets and fall behind once the number of rows grows.

Related event: NeurIPS Review Controversies and the Edge of Native Numerical PFNs(2 posts)→

Original post →

More from Research

Research channel →