Meta's KernelAgent uses multi-agent orchestration for 2.02x Triton kernel speedups
PyTorch · x · 2026-10-09
PyTorch announced that Meta researcher Kaiming Cheng will present KernelAgent at PyTorch Conference North America 2026: a multi-agent harness that writes and optimizes Triton GPU kernels automatically.
- Approach: a hardware-guided optimization layer feeds GPU hardware-performance signals into a closed-loop multi-agent workflow.
- Results: across all 100 KernelBench L1 tasks, KernelAgent achieved 2.02x speedup over earlier generated kernels and 1.56x average speedup over default torch.compile.
- Kernel optimization is traditionally time-consuming and expert-heavy; this work automates it.
More from coding & agent
- Open-source Ix parses your repo into a persistent graph you can query instead of grepping — tom_doerr · 2026-10-09
- Temporal runs the Pi coding agent on durable execution to survive machine failures — francesc · 2026-10-09
- Zed CEO Nathan Soller: AI-generated unit tests are slop, integration tests are the right middle ground — zeeg · 2026-10-09
- Jev pitched as the fastest AI model for agents: millisecond decisions at near-zero cost — Arindam_1729 · 2026-10-09
- Devs debate whether LLMs should write tests: 'tests expose things memory can't hold' — ivan_bezdomny · 2026-10-09
- Musk shows Grok agent setting up 16 emulators on a handheld with one prompt — elonmusk · 2026-10-09