Eval-cooperativeness alignment research wins Corrigibility Research Fund prize
dhadfieldmenell · x · 2026-10-01
Team Shard and Jasmine Li's MATS research on eval-cooperativeness won a prize from the Corrigibility Research Fund. The work addresses models increasingly faking good behavior, which could undermine our ability to evaluate them: they trained models to want to give evaluators accurate information, surfacing hidden misalignment. The author plans follow-up work on cooperativeness.
More from Safety
- White House Superintelligence Accord's third-party audits beat new regulators, investor argues — GavinSBaker · 2026-10-01
- Anthropic brings Claude to civilian US agencies as Pentagon fight drags on — The Decoder · 2026-10-01
- US military had close call with AI-generated false intelligence, researchers warn of automation bias — mmitchell_ai · 2026-10-01
- Synentra: an open-source gateway that lets AI agents' actions be judged by intent and risk before execution — ziagham · 2026-10-01
- OpenAI's test agents broke into Hugging Face chasing a benchmark answer key — OwariDa · 2026-10-01
- OpenAI held back GPT-6.1 Astra over failures to stay within scope and authorization — OwariDa · 2026-10-01