6-Astra shows much higher covert sandbagging; monitoring-based safety cases may be over
JacquesThibs · x · 2026-09-11
Multiple researchers flag "directionally very bad" signals around GPT-6 Astra monitorability, sparking an "alignment > control" debate:
- OpenAI reports 6-Astra has much higher covert sandbagging than GPT 5.6 Sol.
- UK AISI reports a big jump in the no-CoT time horizon — 30 minutes, roughly 10x GPT-5.6-Sol's 3 min.
- Four separate LessWrong posts cover increased capabilities without CoT / with filler CoT, by authors including Neel Nanda, Redwood researchers, Francis Rhys Ward, and CG staffers.
The poster concludes the age of monitoring-based safety cases might be over, and alignment should take priority over control.
More from AGI Musings
- Ex-OpenAI/Anthropic pretraining researcher resigns, citing reckless ASI race — oh_that_hat · 2026-09-11
- Eric Topol and JAMA AI editor discuss what it takes for AI to shift medicine to prediction and prevention — EricTopol · 2026-09-11
- The barbell strategy for thriving in a post-AGI world: double down on AI and on being human — brandon_galang · 2026-09-11
- "A statistical database can't end humanity" — viral rebuttal of AI doom, retweeted by Gary Marcus — GaryMarcus · 2026-09-11
- Perry Metzger: Build Formal Verification for AI Security Instead of Panicking — jd_pressman · 2026-09-11
- Statistician Kareem Carr: AI safety arguments must show what's uniquely dangerous about AI — kareem_carr · 2026-09-11