Claude Opus 5.5 writes 35.5x kernel on KernelBench-Mega, beating GPT-6 Astra's 24.8x
scaling01 · x · 2026-09-23
On KernelBench-Mega, Claude Opus 5.5 [xhigh] produced a Kimi-Linear decode megakernel for the RTX PRO 6000 hitting 35.5x the optimized PyTorch baseline, with GPT-6 Astra second at 24.8x. A performance engineer running 20 instances in parallel says the reported 'nerf' may only apply to TPU/Trainium rather than CUDA, and praises the model's writing style and agency for letting him manage many parallel runs with scarce human attention, calling it a massive jump over fable and astra.
More from coding & agent
- Moda's agent observability weekly: Jev for sharper signals, whole-conversation analysis for long-running agents — KlausCodes · 2026-09-23
- Opus 5.5 builds a Lanterns Festival scene in one prompt, using just 13% of a weekly Claude Max budget — Silver-Chipmunk7744 · 2026-09-23
- stuntd: a local proxy that learns your LLM decisions and serves them at 22ms without an API key — Inevitable-Log5414 · 2026-09-23
- Metriqual: an infra layer that lets AI agents survive model outages by persisting their state — its_vayishu · 2026-09-23
- Adding OAuth to a Sonos MCP server: 6-digit codes beat redirects, discovery metadata matters most — Mean-Gazelle5347 · 2026-09-23
- Why are there no anti-slop coding evals? 70% on frontierSWE but 100 BS tests — yacineMTB · 2026-09-23