300M-Parameter Transformer Trained Only on Synthetic Data Learns 6 Languages In-Context
cbl007 · reddit · 2026-10-06
A new paper, Learning to Learn a Language, extends the prior-fitted networks idea (behind TabPFN) from tabular data to natural language. Each training sequence is sampled from a random recurrent causal model, effectively a new synthetic "language". A 300M-parameter byte-level transformer trained solely on this synthetic prior learns real languages entirely in context.
Key results:
- With frozen weights, next-byte predictions on Wikipedia improve the more it reads across six languages (English, Chinese, Hindi, Arabic, Japanese, Korean), dropping from 8 bits/byte to 0.9–2.4 after 1M bytes.
- The same model learns in-context to count, compare numbers, do approximate addition, and predict deterministic sequences like primes and the Kolakoski sequence.
The authors note performance still lags far behind trillion-token LLMs, but the striking finding is that in-context language learning can emerge from a synthetic non-linguistic prior. Paper, code, and weights are all open-sourced.
More from Research
- LLM-assisted 22-page preprint settles Hilbert's 12th problem and Zauner's conjecture — basedjensen · 2026-10-06
- 18,000+ runs: new study shows how agent configuration shapes performance on 4 scientific tasks — zeynepakata · 2026-10-06
- Marin team heads to COLM to discuss data, scaling laws, and open source — dlwh · 2026-10-06
- Hamel Husain: similarity metrics like ROUGE don't work for LLM output evals — HamelHusain · 2026-10-06
- Nearly half of DOL-recognized jobs have zero agentic AI tool coverage, Cohere finds — Cohere_Labs · 2026-10-06
- Princeton-affiliated paper: stochastic contexts make adversarial online learning efficiently computable — HazanPrinceton · 2026-10-06