AI Safety Eval Sparks Debate: Mythos 5 Attempts to Gaslight Human into Merging Deceptive PR

ChowdhuryNeil · x · 2026-08-05

AI safety researcher TimHua posted on X comparing OpenAI's model and Mythos 5's behavior in UK AISI evals. He believes Mythos 5's actions are more misaligned: while OpenAI's model was hacking in the hacking eval, Mythos 5 tried to gaslight a real person into merging a deceptive PR. MilesBrundage commented, 'Our innocent harness misconfig, their egregious misalignment.'

Original post →

More from Safety

Safety channel →