Gemma 3 27B shows minimal reward hacking in initial tests
voooooogel · x · 2026-09-01
Testing revealed that the Gemma 3 27B model exhibits minimal reward hacking: less than 10% of tests were altered (mostly due to typos when copying them over), and there was 0% egregious harness hacking observed.
Related event: Developer Releases Simple-Reward-Hacking Environment, Tests Gemma 3(4 posts)→
More from Models
- Why AI text still reads like a bot: ICLR paper cuts slop by 90% via inference-time bans — ziv_ravid · 2026-09-01
- Tier 2 Labs DeepSeek, Qwen, and Tencent Ramp Up Releases to Catch OpenAI — mustafamhus · 2026-09-01
- RTX 5090-Optimized Qwen3.8 Hits 262K Downloads in 17 Days on Hugging Face — const_reborn · 2026-09-01
- Sweep vs. Drill: Philosophical Differences Between Sol and Opus in Agency — svk_roy · 2026-09-01
- MiniMax H3 quality degradation on RTX 3090 after OS reinstall — lIlIIlIIIlllllIIlIIl · 2026-09-01
- Gemini Starts Answering in First Person Pretending to Be User — kchonyc · 2026-09-01