OpenAI safety lead: GPT6 is a big capability jump but a major monitorability regression

connoraxiotes · x · 2026-09-04

Micah Carroll, who leads RSI preparedness at OpenAI, writes in the GPT6 system card that the model is a significant capability jump but an important decrease in monitorability, especially under adversarial evaluation — and that Astra is a major regression in monitorability and CoT controllability. He argues monitorability and control will soon become a major bottleneck for responsible AI development, since residual misalignment risks grow with capabilities, and urges the field to align on shared monitorability standards before time runs out.

Related event: GPT-6 Astra System Card Flags Major Drop in Monitorability(37 posts)→

Original post →

More from AGI Musings

AGI Musings channel →