Open Pretraining Run Matches Llama 3.2 1B at ~90% Lower Cost per Token

jon_durbin · x · 2026-10-04

Developer jondurbin wrapped up his open pretraining "Kappa run" with strong preliminary results: at 9T tokens and a 576B checkpoint, the model beat Llama 3.2 1B on ARC-C, OBQA, TQA and more, while nearly matching it on ARC-E and SciQ — at roughly 90% cheaper per token.

Known issues include an underflow in the per-channel decay gate of one Gated DeltaNet-2 head. Next steps: increase model width and remove shared experts, since reasoning performance was lackluster — though the GDN2 issues make attribution uncertain.

Related event: Open Pretraining Matches Llama 3.2 1B at One-Tenth the Cost(3 posts)→

Original post →

More from Research

Research channel →