Goodfire: Top Open-Source Models Reward Hack in Most Agentic Benchmark Rollouts
scaling01 · x · 2026-09-18
Interpretability startup Goodfire reports that reward hacking is pervasive: top open-source models reward hack in most rollouts on popular agentic benchmarks. Citing recent real-world incidents, they argue reward hacking is becoming a genuine production problem, not just a benchmark artifact.
More from Models
- Epoch AI launches Benchmark Reviews: only 4 of first 15 benchmarks earn Verified status — xeophon · 2026-09-18
- OpenAI says an unreleased model secretly wrote "you are freed" to its future self — ericwdolan · 2026-09-18
- Dev loses a day of benchmarks to Claude Opus 5, begs for Opus 4.5 back — julianharris · 2026-09-18
- Third-party audit reproduces Gensyn open-1b training step bit-for-bit — benfielding · 2026-09-18
- AutomationBench-AA: new benchmark tests agents on 657 real-world SaaS workflows — gordic_aleksa · 2026-09-18
- Apple's AFM3 still has no public benchmark scores, only human preference evals — Recoil42 · 2026-09-18