Hand-written Mojo Metal kernels train GPT-2 124M on an M4 Max 1.71x faster than PyTorch MPS
ulmentflam · reddit · 2026-07-22
- The author ported Karpathy’s llm.c to Mojo, added a Metal backend, and trained GPT-2 124M on an M4 Max with no PyTorch or CPython at train time.
- On the benchmarked setup, the Mojo/Metal path reached 8138 tok/s for bf16 and was 1.71× faster than PyTorch MPS bf16; MLX still won overall at 10077 tok/s.
- The post details bottlenecks: matmul dominated step time, Metal bf16 matmul was only about 1.1× faster than fp32, while MLX’s bf16 path used tensor cores for about 2×.
- Correctness was gated with gradient checks, a 10-step loss trajectory, and a 235-test equivalence suite; the repo also includes a FineWeb checkpoint scoring 29.53% on HellaSwag.
- The author says the work was fully hand-written at the kernel/trainer level, and concludes Mojo required more device-specific branching than expected.
More from Infra
- Primis explores hedging infrastructure to make AI compute pricing predictable — dolos_diary · 2026-07-22
- PyTorch Foundation rolls out quarterly updates for six hosted projects — PyTorch · 2026-07-22
- Puri.li opens a free web search API with a 185M-page index — skillplayed · 2026-07-22
- Open-Source Coding Agent Octomind Adds Hosted Machines and 21-Model Access — donk8r · 2026-07-22
- Atome LM beats or matches TFLite Micro on 18 MCU tasks while staying 5× to 70× smaller — themoroccanship · 2026-07-22
- Compute shortage could let Amazon, Microsoft and Google lift margins — RihardJarc · 2026-07-22