AccelEval benchmark: LLMs turn CPU code into GPU kernels up to 161x faster
新智元 · wechat · 2026-10-01
Researchers introduce AccelEval (accepted to NeurIPS 2026), a benchmark testing whether LLMs can turn correct-but-slow CPU programs into GPU implementations that are correct and faster end-to-end. Unlike KernelBench-style operator benchmarks, it spans 6 domains (scientific simulation, finance, graph algorithms, operations research, etc.), 42 tasks, 3 input sizes, and stresses that a faster kernel doesn't guarantee a faster program once data transfer and teardown overheads are counted.
On H200, Gemini 3.1 Pro achieved a 71.2x geometric-mean speedup over a single-threaded CPU baseline on mid-size validated tasks (161.5x at large scale), yet the best model output was still only 0.75x as fast as human CUDA references on tasks with one. The team also built a taxonomy of 43 CUDA optimization strategies, finding structural optimizations (kernel fusion, tiling) correlate with performance, and task-specific strategy guidance yields 1.78x improvement versus 1.18x for generic advice.
More from Infra
- Broadcom to lend Anthropic up to $42B for AI chips, eyeing top customer slot by 2027 — rohanpaul_ai · 2026-10-02
- Meta paper: only 50-60% of recommendation training time actually trained before optimizations — _reachsumit · 2026-10-02
- Dev open-sources GPT-2-tools to run original 1.5B GPT-2 XL locally on CPU — MikePFrank · 2026-10-02
- CoreWeave launches serverless GPUs: hourly-billed, no contract, private preview — altryne · 2026-10-02
- RTX 5090 + 5070 Ti workstation: doubling VRAM wasn't worth it for local LLMs — Lordofwhut · 2026-10-02
- 6x BC-250 mining board cluster runs local LLMs at 100k context, 28 tok/s — Ok-Breadfruit-3523 · 2026-10-02