Mutation Analysis Shows GPU-Kernel Benchmarks Miss 16.9% of Faults

NationalUniversityofSingapore · hf · 2026-09-22

Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and loose float tolerances, and their verdicts feed leaderboards and RL rewards. This work introduces mutation analysis as an adequacy metric for kernel-benchmark oracles: 10,303 compilable faults injected into verified CUDA implementations of 188 KernelBench problems (7,384 with independent kill witnesses), scoring any test protocol by detection rate.

Key findings:

Released as the KernelBench-M dataset.

Original post →

More from Research

Research channel →