NanoGPT speedrun record falls to 39.9s, -46% via flop-level skipping tricks

yacinelearning · x · 2026-09-29

modded-nanogpt has a new record: devenpzak (MIT Math) cut training time from 67.6s to 39.9s (-34.0s, -46% on the same hardware); PR #360 was merged and independently reproduced by four sources. The new paradigm: optimize at the individual flop level rather than just matmuls or fancier ops—skip low-value flops entirely. Key tricks include sampled softmax (8s saved, skipping lmhead fwd/bwd for tokens absent from a batch), sparse optimizer steps only for ngram embeddings that occurred in the batch (beta1=0, beta2 applied retroactively), full-stack fp8, a new embedding table, and ANVIL2.

Original post →

More from Research

Research channel →