Open pretraining run matches Llama 3.2 1B on ARC-C at ~10% of the cost, author details 4 pitfalls
jon_durbin · x · 2026-10-03
jondurbin shared preliminary results and a post-mortem of the Kappa pretraining run:
Results: At the 576B-token checkpoint, the model beat Llama 3.2 1B (trained on 9T tokens) on ARC-C, OBQA, TQA and others, and nearly matched it on ARC-E, SciQ, etc. — at roughly 90% lower cost per token.
Issues encountered:
- A per-channel decay gate in one Gated DeltaNet-2 head underflowed, wiping state; output RMSNorm then amplified near-zero outputs 1000x into fleet-wide grad spikes;
- Very slow/late nodes had their evidence discarded; better tolerance for stale data and partial updates is needed;
- Samples were packed without masking, suboptimal for pretraining;
- GDN2 in the first and last layers may be suboptimal vs. SWA, though evidence is unclear.
He calls it "a solid preliminary win" and plans fixes.
Related event: Open Pretraining Run Matches Llama 3.2 1B at One-Tenth the Cost(2 posts)→
More from Infra
- Modal VM Sandboxes hit GA: demo runs Docker Compose apps, tests and coding agents — charles_irl · 2026-10-03
- Community poll of 793 MLX users: oMLX wins at 55.5%, dominating Ultra chips — HankYeomans · 2026-10-03
- Local AI user asks: why does everyone enjoy taming the beast of complexity? — kathi7 · 2026-10-03
- AI gateways emerge as key control layer, with eight core capabilities to govern production AI stacks — goyalshaliniuk · 2026-10-03
- 8GB VRAM folks' daily prayer for Qwen4 35B A3B, settling for Gemma 26B QAT — RobustLokiX · 2026-10-03
- Two 300B MoE models on one 128GB Strix Halo: Kyojin engine hits 44 tok/s decode — Yaniss916 · 2026-10-03