PFLM: a 300M model pretrained on zero real languages learns languages in context

Lennart Carstens-Behrens · hf · 2026-10-06

PFLM (Prior-Fitted Language Model) is a 300M-parameter byte-level transformer pretrained only on sequences from a synthetic non-linguistic prior—recurrent structural causal models drawn fresh from a distribution, with no real language ever seen and no language repeated. Frozen weights, pure in-context learning: the only way to predict a continuation is to infer the language from its prefix.

The model hasn't learned a language—it has learned to learn one.

Original post →

More from Research

Research channel →