ANVIL III Optimizer Claims 62% Pretraining Cost Cut at Frontier Scale, Beats Tuned Muon

kellerjordan0 · x · 2026-10-07

Hyperstition introduces ANVIL III, calling it the biggest optimizer breakthrough since Muon: 62% pretraining cost reduction at an 8x Chinchilla budget, 30% lower inference costs via faster decode, and 20–28 millinats better final loss than fully tuned Muon at the same compute. Efficiency gains scale from 32% at 124M to 50% at 1.2B params. Their Feather 1.7B model beats Qwen3-1.7B-Base on math with 180x fewer training tokens, and the team holds the NanoGPT speedrun record via ANVIL II. The post traces the Shampoo→Muon matrix-preconditioned lineage and clarifies the speedrun was an efficiency challenge, not a scaling effort.

Original post →

More from Infra

Infra channel →