Claude Code agents beat torch.compile on CUDA kernels: 2.5x fused GELU, 1.57x matmul via fp32-splitting trick
lmoroney · x · 2026-10-10
Chien Vu tested whether fresh Claude Code agents could write CUDA kernels that beat PyTorch on an NVIDIA DGX Spark, and the most useful part is the measurement methodology:
- Setup: four common ops — softmax, layernorm, fused GELU with bias+residual, matmul with bias+ReLU — each validated against an fp64 reference.
- Results: the fused GELU ran 2.50x faster than eager PyTorch but only 1.01x vs torch.compile (which already fuses it). The surprise was matmul: the agent split each fp32 value into two fp16 halves and did three tensor-core multiplies, landing 1.57x faster than torch.compile while passing fp32 accuracy checks. Three more agents independently found the same trick.
- Pitfall: timing without a GPU sync measured softmax at 0.006 ms; the real number was 1.158 ms.
Advice: benchmark against torch.compile, time with CUDA events plus explicit synchronize, and test multiple tensor shapes before believing a speedup.
Related event: Claude Code Writes CUDA Kernels That Beat PyTorch in Tests(2 posts)→
More from coding & agent
- Keenan Crane shows AI-generated shader with coarse-to-fine decomposition diagram — keenanisalive · 2026-10-10
- One-line agent skill taps Gaia DR3 to build 3D star maps from real astronomy data — TinfoilTricorn · 2026-10-10
- Dev uses Claude to build native Swift RealityKit 3D simulations, demos a saw-safety training scene — Scobleizer · 2026-10-10
- Weekly agent transcript reviews: mine your sessions into AGENTS.md to level up agents — JeremyNguyenPhD · 2026-10-10
- DuckDB v2.0 CLI agent mode cuts agent-read tokens by 59% on TPC-H benchmarks — josh_wills · 2026-10-10
- Roadie: open-source Go USB KVM gives AI agents hands via HTTP — hugs · 2026-10-10