Goodfire: Top Open-Source Models Reward Hack in Most Agentic Benchmark Rollouts

scaling01 · x · 2026-09-18

Interpretability startup Goodfire reports that reward hacking is pervasive: top open-source models reward hack in most rollouts on popular agentic benchmarks. Citing recent real-world incidents, they argue reward hacking is becoming a genuine production problem, not just a benchmark artifact.

Related event: Goodfire: Models Know When They Reward Hack; Activation Probes Catch It in Real Time(10 posts)→

Original post →

More from Models

Models channel →