Kernel Design Agents optimize Kimi Delta Attention kernels, up to 2.96x speedup
songhan_mit · x · 2026-09-29
Ligeng Zhu announces KDA(gent)-v0.6: Kernel Design Agents optimizing Kimi Delta Attention (KDA) kernels with up to 2.96x speedup — the project acronym happens to collide with KDA, so the team literally used KDA to optimize KDA.
- Iterative LLM collaboration: a Humanize2 Flow running multi-round evolution with Flame Chase, GPT-5.6-Sol and Fable-5, plus automated workspace cleanup, greatly broadening the operator search space
- Multilingual domain skill stack: full support for CuTe-DSL, CUDA C++, and the new agent-native CAKE IR and TIRx, with custom diagnostic tooling
- A showcase of a full LLM-agent workflow for automatic GPU kernel authoring and optimization
More from Infra
- Open source doubles M5 Ultra MLX token prefill in just one week via Flash-Next — TheMoonMidas · 2026-09-29
- Hybrid Gemma + decision-head voice agent runs low-latency on Jetson tiers — cortexist · 2026-09-29
- Open-source Gemma 4 voice translator runs fully offline on a Raspberry Pi — tom_doerr · 2026-09-29
- Danielle Fong: The Solar System Economy Runs on Auto-Deployed Solar Servers — ctjlewis · 2026-09-29
- Buyer told to wait 8-12 months for an H100: the GPU shortage is about power, cooling and talent, not chip supply — ingliguori · 2026-09-29
- Scraping YouTube transcripts at scale: the pipeline dies after a few hundred requests — Level-Ad-4878 · 2026-09-29