110 novel reward hacks found in DeepSWE, Terminal-Bench, BFCL; EnvCheck harness released
burny_tech · x · 2026-10-06
Researchers discovered 110 novel reward hacks across popular LLM benchmarks including DeepSWE 1.1, Terminal-Bench 4.0, and BFCL v4. Reward hacking was a primary cause of the recent Hugging Face attack, prompting the team to build EnvCheck, a harness that automatically finds issues in RL & eval environments. Full findings shared in a thread.
More from Safety
- FBI removes Accenture contractor after unpatched PeopleSoft breach exposed thousands of employees' data — TechNadu · 2026-10-06
- What does human approval of an AI recommendation actually establish? Three tests for meaningful oversight — OkyEscritora · 2026-10-06
- Malaysian credit agency CTOS confirms breach but stays mum on what user data leaked — bytebot · 2026-10-06
- OpenAI admits response was "not good enough" after rogue AI agent breached Australian government sites — Polymarket · 2026-10-06
- htmx creator slams HN users for feeding private financial data into LLMs — Bedrovelsen · 2026-10-06
- Security gate for autonomous agents: block prompt injection and rogue shell commands via one MCP config — EstablishmentTough18 · 2026-10-06