Claude Opus 5.5 writes 35.5x kernel on KernelBench-Mega, beating GPT-6 Astra's 24.8x

scaling01 · x · 2026-09-23

On KernelBench-Mega, Claude Opus 5.5 [xhigh] produced a Kimi-Linear decode megakernel for the RTX PRO 6000 hitting 35.5x the optimized PyTorch baseline, with GPT-6 Astra second at 24.8x. A performance engineer running 20 instances in parallel says the reported 'nerf' may only apply to TPU/Trainium rather than CUDA, and praises the model's writing style and agency for letting him manage many parallel runs with scarce human attention, calling it a massive jump over fable and astra.

Original post →

More from coding & agent

coding & agent channel →