NanoGPT speedrun record falls to 39.9s, -46% via flop-level skipping tricks
yacinelearning · x · 2026-09-29
modded-nanogpt has a new record: devenpzak (MIT Math) cut training time from 67.6s to 39.9s (-34.0s, -46% on the same hardware); PR #360 was merged and independently reproduced by four sources. The new paradigm: optimize at the individual flop level rather than just matmuls or fancier ops—skip low-value flops entirely. Key tricks include sampled softmax (8s saved, skipping lmhead fwd/bwd for tokens absent from a batch), sparse optimizer steps only for ngram embeddings that occurred in the batch (beta1=0, beta2 applied retroactively), full-stack fp8, a new embedding table, and ANVIL2.
More from Research
- ParaSpeechCLAP Nominated for Best Student Paper at Interspeech 2026 — jessyjli · 2026-09-29
- Physical modeling of Micro-C data reveals euchromatin forms condensed surface-active domains — anshulkundaje · 2026-09-29
- BOBseq preprint on bioRxiv brings RNA isoform resolution to scalable transcriptomics — anshulkundaje · 2026-09-29
- Matryoshka Attribution: reverting 1% of weights removes most refusal in Llama 3.1 8B — aryaman2020 · 2026-09-29
- Romero Lab builds Chat with PyMOL, letting researchers talk to molecular structures in plain language — rbhar90 · 2026-09-29
- Paper2Agent turns research papers into interactive AI agents, Nature paper shows — james_y_zou · 2026-09-29