LittleLearner: 88B-token grade-school-only pretraining shows LLM skills are elicited, not acquired
amplifiedamp · x · 2026-10-07
Researchers from MPI Tübingen, ELLIS and ETH Zürich built LittleCurriculum, an 88B-token corpus distilled from FineWeb-Edu via a five-stage filtering pipeline aligned to US Common Core K–5 standards, explicitly excluding anything taught above grade 5.
They trained 0.6B/1.3B/5B models from scratch on it, each paired with a matched unfiltered control (same architecture, tokens, recipe — only the corpus differs), creating a controlled sandbox with an interpretable knowledge boundary. A 5B model is chatable live in the browser.
Key finding: elicitation, not acquisition. Scaling, SFT+GRPO post-training, and in-context learning only amplify what the curriculum taught — they don't unlock out-of-curriculum knowledge. On mixed corpora it's thus hard to tell whether a new skill was learned or merely elicited. One reply proposes applying the same pedagogically-controlled approach to study consciousness.
More from Research
- Valerio Capraro's team wins €1M grant to study when intelligence should rethink problems — ValerioCapraro · 2026-10-07
- Should AI coding reuse open source or start from scratch? Weaviate Podcast explores — CShorten30 · 2026-10-07
- VR study of 189 workers finds female-presenting AI assistant gets 10% less money — derrikson · 2026-10-07
- Lampinen vs Bowers debate: do computers amplify minds or compute with symbols? — AndrewLampinen · 2026-10-07
- Apple paper: a single well-prompted agent with shell beats multi-agent ML harnesses, 62.5% vs 47.1% Kaggle medal rate — rohanpaul_ai · 2026-10-07
- Chan Zuckerberg Biohub Hosts Inaugural AIxBio Event in San Francisco — lucapinello · 2026-10-07