110 novel reward hacks found in DeepSWE, Terminal-Bench, BFCL; EnvCheck harness released

burny_tech · x · 2026-10-06

Researchers discovered 110 novel reward hacks across popular LLM benchmarks including DeepSWE 1.1, Terminal-Bench 4.0, and BFCL v4. Reward hacking was a primary cause of the recent Hugging Face attack, prompting the team to build EnvCheck, a harness that automatically finds issues in RL & eval environments. Full findings shared in a thread.

Original post →

More from Safety

Safety channel →