AccelEval benchmark: LLMs turn CPU code into GPU kernels up to 161x faster

新智元 · wechat · 2026-10-01

Researchers introduce AccelEval (accepted to NeurIPS 2026), a benchmark testing whether LLMs can turn correct-but-slow CPU programs into GPU implementations that are correct and faster end-to-end. Unlike KernelBench-style operator benchmarks, it spans 6 domains (scientific simulation, finance, graph algorithms, operations research, etc.), 42 tasks, 3 input sizes, and stresses that a faster kernel doesn't guarantee a faster program once data transfer and teardown overheads are counted.

On H200, Gemini 3.1 Pro achieved a 71.2x geometric-mean speedup over a single-threaded CPU baseline on mid-size validated tasks (161.5x at large scale), yet the best model output was still only 0.75x as fast as human CUDA references on tasks with one. The team also built a taxonomy of 43 CUDA optimization strategies, finding structural optimizations (kernel fusion, tiling) correlate with performance, and task-specific strategy guidance yields 1.78x improvement versus 1.18x for generic advice.

Original post →

More from Infra

Infra channel →