AI agent ports and optimizes CUDA attention kernel to CuTeDSL in an afternoon, 1.2x faster
knowrohit07 · x · 2026-09-17
In a new article, maharshii reports his AI agent ported and optimized a CUDA C++ attention kernel to CuTeDSL in a single afternoon — the result ran 1.2x faster than the original kernel and 2.8x faster than cuDNN.
Commenter knowrohit07 adds that getting comfortable with tensor dialects (lower-level than DSLs or C++) pays off: he has written kernels directly in MLIR and ported parts of the CuTe DSL, arguing DSLs generally have poor debugging experiences.
More from coding & agent
- Let Codex audit your skills and agents.md to fit GPT-6 Astra — nickbaumann_ · 2026-09-17
- Stealth model Union Alpha scores 74% on DeepSWE, beating GPT-5.6 Sol at lower cost — ZeroStateReflex · 2026-09-17
- Agentic AI systems are the next network users: 40% of enterprise apps to include agents by 2026 — seankinneyRCR · 2026-09-17
- 2012 iPad mini runs 90k-line Zig game at 60fps with ~100 lines of patches — banteg · 2026-09-17
- xAI launches Grok Build coding agent powered by Grok 4.6 with cross-session memory — ns123abc · 2026-09-17
- Yandex Research: KV-cache as a runtime for concurrent LLM interaction without retraining — RichmanRonald · 2026-09-17