Open Pretraining Run Matches Llama 3.2 1B at ~90% Lower Cost per Token
jon_durbin · x · 2026-10-04
Developer jondurbin wrapped up his open pretraining "Kappa run" with strong preliminary results: at 9T tokens and a 576B checkpoint, the model beat Llama 3.2 1B on ARC-C, OBQA, TQA and more, while nearly matching it on ARC-E and SciQ — at roughly 90% cheaper per token.
Known issues include an underflow in the per-channel decay gate of one Gated DeltaNet-2 head. Next steps: increase model width and remove shared experts, since reasoning performance was lackluster — though the GDN2 issues make attribution uncertain.
Related event: Open Pretraining Matches Llama 3.2 1B at One-Tenth the Cost(3 posts)→
More from Research
- Just-significant economics results at p=0.05 have only ~25% replication odds, new I4R paper — RexDouglass · 2026-10-04
- Research agent claims irrationality measure bound for π of 6.0446, pending expert review — Michael_D_Moor · 2026-10-04
- Bittensor's Synth runs 200+ models forecasting crypto volatility; Volatility CRPS scoring goes live — bittingthembits · 2026-10-04
- RL environments are data: they push the frontier until they saturate — Shahules786 · 2026-10-04
- Josh Gans' new paper argues guilt explains why we demand more from machines than people — joshgans · 2026-10-04
- Study of 18 LLMs Finds No Single Model Excels at Diverse Open-Ended Generation — gregd_nlp · 2026-10-04