MuonH with a minus-sqrt LR schedule cuts modded-nanogpt Track 3 by 75 steps
liliang_ren · x · 2026-07-21
We submit MuonH with a minus-sqrt learning-rate schedule \((1-\sqrt{t})\) to the modded-nanogpt Track 3 benchmark.
- The new schedule improves the current MuonH SOTA by 75 steps.
- It also nearly doubles the peak learning rate under this setup.
- The comparison chart shows the new curve converging faster in the tail than the prior MuonH runs.
More results are promised later.
More from Research
- Swaayatt demos autonomous driving at 52 km/h on mountain roads, self-recovers after skid — sanjeevs_iitr · 2026-09-11
- Fast ViT shows strong ImageNet results; scaling runs needed next — ducha_aiki · 2026-09-11
- Loss Functions Are Scientific Assumptions: MSE Implies Gaussian Noise, Cross-Entropy Implies Bernoulli — bravo_abad · 2026-09-11
- SymKit MCP: 44 tools for AI agents to verify symbolic derivations — Foreign-Specific-604 · 2026-09-11
- Researchers: LLMs under pressure invent new languages unreadable to humans — mikeflache · 2026-09-11
- Mi-Ripple fixes ripple artifacts left by iterative AI image editing — Miyang-AI · 2026-09-11