232x faster kernel with Codex auto-research: a GPU Mode contest postmortem
dejavucoder · x · 2026-10-06
A postmortem of the GPU Mode auto-research kernel contest: the author used Codex for "auto-research" (aka loop engineering) to achieve a 232x speedup over baseline on batched Householder QR factorization (qrv2), placing 12th out of 183.
Key takeaways:
- Inference optimization has moved up the abstraction ladder — it's less about hand-writing the fastest kernel and more about defining optimization context, I/O, constraints, and testing environments for your AI agents.
- Learn enough math yourself to ask better questions before handing off to agents.
- Introduce idea diversity to escape local maxima, plus implementation-level prompting hints for breakthroughs.
This was the author's first serious auto-research attempt; bottlenecks and improvements are documented in detail.
More from coding & agent
- Building a RAG-powered FAQ bot: RRF-fused dedup keeps the knowledge base clean — Al_Grigor · 2026-10-06
- A regex bug let 500K rows slip past an LLM SQL guard: two fail-closed checks for multi-tenant warehouses — Hrolgarr · 2026-10-06
- iPad handwriting app SuperrPaper ships MCP server: Claude builds notebooks, you write by hand — AbrocomaMaleficent88 · 2026-10-06
- Dozens of AI agents ran 24/7 with little steering while founder handled admin — kevinnbass · 2026-10-06
- 5 principles for building AI memory: store, structure, retrieve, update, forget — goyalshaliniuk · 2026-10-06
- Autoresearch loops 'two months away' sparks debate over hyperparameter tuning line — A_K_Nain · 2026-10-06