Limite trained from scratch on under 300B math-heavy tokens, nanogpt-speedrun-inspired
tensorqt · x · 2026-09-22
Paradigm reveals that Limite was pretrained from scratch on fewer than 300B tokens — almost entirely math — using an architecture inspired by recent advances from the nanogpt speedrun competitions. A strikingly small, narrow corpus for a model targeting hard mathematics.
More from Models
- Gemini beats GPT-6 Astra at robot capture the flag, winning 70% of matches — chris_j_paxton · 2026-09-22
- 105 Planted Bugs Put Grok 4.7 at 28.7 vs GPT-6 Astra's 45 in Real-Repo Coding Test — PawelHuryn · 2026-09-22
- Regression on OpenAI's GPT-5.6 Benchmark Data Backs Out Their λ Values — tobyordoxford · 2026-09-22
- Developer Burns 1.46B Tokens in 13 Hours, Jokes He's "Part of the Infrastructure" — MaziyarPanahi · 2026-09-22
- Swarm scaling needs squared inference to match chain-of-thought gains, analysis of OpenAI curves finds — tobyordoxford · 2026-09-22
- User claims Gemini 4 Pro in production isn't the model on benchmarks — Ambroverse · 2026-09-22