232x faster kernel with Codex auto-research: a GPU Mode contest postmortem

dejavucoder · x · 2026-10-06

A postmortem of the GPU Mode auto-research kernel contest: the author used Codex for "auto-research" (aka loop engineering) to achieve a 232x speedup over baseline on batched Householder QR factorization (qrv2), placing 12th out of 183.

Key takeaways:

This was the author's first serious auto-research attempt; bottlenecks and improvements are documented in detail.

Original post →

More from coding & agent

coding & agent channel →