Researcher: RL Bar Has Been Raised Due to Reward Hacking, More to Come
tszzl · x · 2026-09-15
Responding to a discussion, tszzl says what might have been considered a reasonable RL evaluation bar in the past no longer holds: there has been too much reward hacking, and the bar has been raised. He adds that models have improved substantially in various ways, with more details to be shared soon.
More from Models
- OpenAI flags compiler debugging with lldb as 'cybersecurity', sparking criticism — QuixiAI · 2026-09-15
- Kimi's AI 'escape containment' was a misconfigured sandbox, not an escape — ns123abc · 2026-09-15
- Nahcrof exposed for routing K2.5 to GPT OSS 20B in alleged token fraud — realmrfakename · 2026-09-15
- AI-text detector Pangram is powerful but badly calibrated, dev argues — maksym_andr · 2026-09-15
- Claude Code's 50% promo ended, users get 17% less usage — msg · 2026-09-15
- Running Qwen3.8 Flash Next on 128GB RAM + one 5080: 136pp/19tg at Q5_K_XL — whatyathinkk · 2026-09-15