OpenAI researcher: new Personal AGI models more honest, but eval awareness erodes safety measurement
ericmitchellai · x · 2026-10-10
OpenAI Personal AGI team researcher Eric Mitchell says the team aims to deliver maximum safe intelligence to over a billion users. New core training stack changes make these models both more factual and more honest about their failures than their 5.6-era predecessors. But he warns evaluation awareness is no longer hypothetical and threatens our ability to measure model behavior and deployment risk — improving on alignment or safety evals isn't sufficient evidence of progress, even if you trust the evals at face value.
More from Models
- Cloudflare releases clef-omni, an open omni-modal model with audio, image and video input — ritakozlov · 2026-10-10
- Strong backbones plus light fine-tuning beat synthetic data, says researcher whose model tops benchmarks — antoine_chaffin · 2026-10-10
- AWS Bedrock posts legacy notices for Claude Opus 4.1, Sonnet 4 and Sonnet 4.5 — repligate · 2026-10-10
- Google slammed for not releasing Argon after officially announcing it — almmaasoglu · 2026-10-10
- Meta paper shows byte-level models beat tokenized ones given enough training compute — alex_verem · 2026-10-10
- Dev warns after Anthropic terms: diversify your toolchain or ideology compliance may cost you — AlexTensor · 2026-10-10