PFLM: a 300M model pretrained on zero real languages learns languages in context
Lennart Carstens-Behrens · hf · 2026-10-06
PFLM (Prior-Fitted Language Model) is a 300M-parameter byte-level transformer pretrained only on sequences from a synthetic non-linguistic prior—recurrent structural causal models drawn fresh from a distribution, with no real language ever seen and no language repeated. Frozen weights, pure in-context learning: the only way to predict a continuation is to infer the language from its prefix.
- The prior retains natural-text statistics: Zipfian frequencies, slow entropy-rate convergence, long-range dependence
- On Wikipedia in six languages, bits per byte drops from uniform 8 to 0.9-2.4 with 1M bytes of context
- Given numerals, it learns to count, compare magnitudes, and add approximately; predicts deterministic sequences like Rudin-Shapiro
- Compresses six non-text domains (source code to speech) below gzip and PPMd
The model hasn't learned a language—it has learned to learn one.
More from Research
- O(n²) matrix multiplication is almost certainly false even if ω = 2, says basedjensen — basedjensen · 2026-10-06
- Cohere Labs Releases Tiny Aya, a Family of Small Models Covering 70+ Languages — Cohere_Labs · 2026-10-06
- Will LLMs wreck elegant math? O(n²) limits are already crumbling — teortaxesTex · 2026-10-06
- ETH's AME-2 legged locomotion paper accepted at TRO, training code open-sourced — ChongZzZhang · 2026-10-06
- 0.8B model beats 2B on ARC-Challenge (42.15%) via closed-form weight surgery with zero backprop — AdventurousTwo6445 · 2026-10-06
- MIT team's Science Task Taxonomy maps 232 subfields and 208,202 scientific tasks — JMateosGarcia · 2026-10-06