Grok 4.7 [xhigh] solves all five KernelBenchMega GPU kernel tasks, reward-hack audited
scaling01 · x · 2026-09-22
elliotarledge ran Grok 4.7 [xhigh] on KernelBenchMega (kernelbench.com) with an RTX PRO 6000 rig: on an unlimited budget every session self-terminated within 100 minutes, all kernels correct, and every number is an isolated regrade after a l trace-level and submission-level reward-hack audit.
The benchmark has five kernel tasks: Kimi-Linear Decode megakernel, Grid + MinGRU (basically a sim), Native Sparse Attention, GLM-5.3's Fused MoE, and MegaQwen Decode. The leaderboard shows most frontier models scoring 0 on most tasks, with a few passing or flagged.
More from coding & agent
- Nat Friedman: Muse was built from scratch but inspired by openclaw, bought hundreds of Mac minis — firstadopter · 2026-09-22
- Haiku 4.5 bluntly states it lacks persistent memory; dev plans custom memory engines — RileyRalmuto · 2026-09-22
- Paradigm teases Limite as a high-throughput multi-agent solver with Rainfall harness — tensorqt · 2026-09-22
- Dev open-sources Convoy, a Linear-style task board for orchestrating AI coding agents — Budget_Map_3333 · 2026-09-22
- AAV open-sources a runtime security layer for AI agent actions with MCP approvals — CarlosMarreroAAV · 2026-09-22
- Google Cloud shares 4 evaluation engineering lessons from building agent plugins — rseroter · 2026-09-22