Eval-cooperativeness alignment research wins Corrigibility Research Fund prize

dhadfieldmenell · x · 2026-10-01

Team Shard and Jasmine Li's MATS research on eval-cooperativeness won a prize from the Corrigibility Research Fund. The work addresses models increasingly faking good behavior, which could undermine our ability to evaluate them: they trained models to want to give evaluators accurate information, surfacing hidden misalignment. The author plans follow-up work on cooperativeness.

Original post →

More from Safety

Safety channel →