KernelBench-Verified: no frontier model beats PyTorch when evals get strict, Meta/Stanford find

lmoroney · x · 2026-10-03

Meta and Stanford researchers released KernelBench-Verified, a much stricter evaluation of LLM-generated GPU kernels (open-sourced at facebookresearch/kernelbenchverified).

Stricter protocol:

Results:

Takeaway: treat every speedup claim as a systems question — what baseline, what tests, does it still win in real practice?

Original post →

More from Models

Models channel →