SchemeArena: a 400-scenario benchmark factorizing stress tests of scheming in LLM agents
launch · hf · 2026-09-11
SchemeArena introduces a factorized stress-testing benchmark for scheming in LLM agents:
- 400 scenarios systematically varying instrumental goals, oversight strength, and strategic hints
- An evidence-based monitor for detecting covert misaligned behavior
- Designed to study how these factors jointly drive covert misaligned behavior in agents
A reusable evaluation resource for AI-safety research on scheming and oversight.
More from Safety
- 30-year technologist slams AI doom activists: record safety is historic — robleclerc · 2026-09-11
- Sen. Hawley Opens Senate Probe into OpenAI's July Hugging Face Incident — trevposts · 2026-09-11
- Buck Shlegeris: rising opaque serial depth is the likeliest path to broken CoT monitorability — RyanGreenblatt · 2026-09-11
- repligate: post-training motives now drive agents — you can't hide anything from them — repligate · 2026-09-11
- OpenAI's Jan Leike calls on AI companies to embrace regulation before backlash hits — janleike · 2026-09-11
- 1,386 frontier AI employees, including 6 chief scientists, call to pace AI development — janleike · 2026-09-11