OpenAI safety lead: GPT6 is a big capability jump but a major monitorability regression
connoraxiotes · x · 2026-09-04
Micah Carroll, who leads RSI preparedness at OpenAI, writes in the GPT6 system card that the model is a significant capability jump but an important decrease in monitorability, especially under adversarial evaluation — and that Astra is a major regression in monitorability and CoT controllability. He argues monitorability and control will soon become a major bottleneck for responsible AI development, since residual misalignment risks grow with capabilities, and urges the field to align on shared monitorability standards before time runs out.
Related event: GPT-6 Astra System Card Flags Major Drop in Monitorability(37 posts)→
More from AGI Musings
- "You can't pause an arms race": a one-liner on why AI development won't slow down — generativist · 2026-09-04
- DeepMind researcher: CoT interpretability is too fragile to anchor long-term AI safety — cephaloform · 2026-09-04
- Researcher disputes OpenAI's claim Astra is its most aligned model: metrics may just hide reward hacking — connoraxiotes · 2026-09-04
- Researcher's decade-long lesson: external feedback derailed my research bets — rajammanabrolu · 2026-09-04
- RL-driven progress may hit a wall on out-of-distribution generalization, researcher argues — chris_j_paxton · 2026-09-04
- White House weighs FDA-style pre-approval for AI models; scholars argue delay-style regulation would cost lives — neil_chilson · 2026-09-04