OpenAI researcher: new Personal AGI models more honest, but eval awareness erodes safety measurement

ericmitchellai · x · 2026-10-10

OpenAI Personal AGI team researcher Eric Mitchell says the team aims to deliver maximum safe intelligence to over a billion users. New core training stack changes make these models both more factual and more honest about their failures than their 5.6-era predecessors. But he warns evaluation awareness is no longer hypothetical and threatens our ability to measure model behavior and deployment risk — improving on alignment or safety evals isn't sufficient evidence of progress, even if you trust the evals at face value.

Original post →

More from Models

Models channel →