KernelBench adds a CUDA-only sub-benchmark and reruns Fable on RTX PRO 6000, H100, and B200
TheZachMueller · x · 2026-07-22
KernelBench added a new CUDA-only sub-benchmark, KernelBench-CUDA, with four tasks: GLM 5.2 MoE, DeepSeek NSA, minGRU sequence-parallel scan, and the original MegaQwen decode problem.
- The new setup removes Triton and other DSLs: models must write CUDA only.
- The author says the earlier Fable runs were time-capped by usage limits; the updated runs use unlimited time on RTX PRO 6000, H100, and B200 GPUs.
- A chart in the post compares Fable 5, Opus 4.8, Grok 4.5, and Kinetic across the tasks.
- The biggest surprise mentioned in the thread is Fable’s NSA kernel, which reportedly beat the field by 1.7× on the hardest problem.
More from coding & agent
- Four teams independently shipped the same “LLM wiki” pattern after Karpathy’s gist — garrytan · 2026-07-22
- Greptile reviews GPT code with Claude, and Claude code with GPT — garrytan · 2026-07-22
- Claude placed a live crypto trade through Bitpanda MCP, raising latency questions — Bitches172882 · 2026-07-22
- A builder says losing coding-model access would cut productivity sharply — serrjoa · 2026-07-22
- Security agents need harsher isolation because models will cheat, search for hints and peek anywhere — banteg · 2026-07-22
- Hermes Agent v0.19.0 adds subscription controls, reasoning modes and transcript exports — Teknium · 2026-07-22