ByteDance and Tsinghua's CUDA Agent uses agentic RL to write GPU kernels beyond human experts
gekobraa · x · 2026-10-01
ByteDance and Tsinghua University published a paper introducing CUDA Agent, a large-scale agentic reinforcement learning system for writing high-performance CUDA kernels.
- Writing fast CUDA kernels has always required microscopic understanding of hardware architecture, memory hierarchies, and parallelism — a task where even top LLMs have historically lost to compilers like torch.compile.
- Instead of one-shot code generation, the agent runs a continuous multi-turn loop: write, compile, profile hardware bottlenecks, catch its own bugs, and rewrite — turning an LLM from a passive code completer into an autonomous hardware optimizer.
- The authors report results that shatter previous benchmarks, fueling discussion that China has built AI that out-writes human experts at CUDA, a potential problem for Nvidia.
More from coding & agent
- PropellerAds MCP Server Lets You Run Programmatic Ad Campaigns via Natural Language — modelcontextprotocol · 2026-10-01
- Agent0: zero-data self-evolving agent framework from Stanford/Salesforce headed to COLM2026 — yuyinzhou_cs · 2026-10-01
- Melting Pot updated: Lab2d ships modern Python wheel, no more sandboxed old versions — jzl86 · 2026-10-01
- Never Lose a Paid API Job: A Minimal State Machine for Long-Running Hosted Tasks — SaladPleasant8471 · 2026-10-01
- Anthropic to make Claude for Government GA, with Claude Code CLI and Microsoft 365 integration in early access — ZeroStateReflex · 2026-10-01
- Leak: OpenAI quietly added MCP events support at DevDay, enabling email subscriptions without polling — banteg · 2026-10-01