Limite trained from scratch on under 300B math-heavy tokens, nanogpt-speedrun-inspired

tensorqt · x · 2026-09-22

Paradigm reveals that Limite was pretrained from scratch on fewer than 300B tokens — almost entirely math — using an architecture inspired by recent advances from the nanogpt speedrun competitions. A strikingly small, narrow corpus for a model targeting hard mathematics.

Original post →

More from Models

Models channel →