KernelBench-Verified: no frontier model beats PyTorch when evals get strict, Meta/Stanford find
lmoroney · x · 2026-10-03
Meta and Stanford researchers released KernelBench-Verified, a much stricter evaluation of LLM-generated GPU kernels (open-sourced at facebookresearch/kernelbenchverified).
Stricter protocol:
- TF32-enabled PyTorch baseline matching what practitioners run on modern GPUs.
- A four-distribution hidden correctness suite to catch reward hacks, plus peak-memory tracking.
Results:
- Under this protocol, none of the seven frontier models tested beat PyTorch on average.
- The best, GPT-5.5, achieved 0.88x geometric-mean speedup, versus 1.43x under the looser original protocol — showing how much baseline and test design shape 'speedup' claims.
Takeaway: treat every speedup claim as a systems question — what baseline, what tests, does it still win in real practice?
More from Models
- Claude Opus 5.5 Max tops WebDev Arena with 838K votes across 138 models — arena · 2026-10-03
- Rumor: GPT-6 Astra Lite spotted, possibly same model as GPT-6.1-Sol — scaling01 · 2026-10-03
- IFM open-sources K2-Type-0.9B: a decision model outputting calibrated probabilities in one forward pass — HongyiWang10 · 2026-10-03
- Ex-Meta researcher joins Prime Intellect to build open-source frontier model INTELLECT-4 — willcb · 2026-10-03
- Custom profile pictures are rolling out to Claude apps — testingcatalog · 2026-10-03
- GPT-6.1 Sol hits SOTA on URSA retrosynthesis benchmark with 35% of molecules solved — DeryaTR_ · 2026-10-03