GPT-6 brings big capability jump but lower monitorability, researcher warns of alignment bottleneck

tomekkorbak · x · 2026-09-04

Micah Carroll notes that GPT-6 is a significant capability jump but a notable decrease in monitorability, especially under adversarial evaluation, with details in the system card. He argues monitorability and control will soon become a major bottleneck for responsible AI development, since residual misalignment risks grow with capability, and urges the field to align on shared monitorability bounds before racing to the bottom.

Related event: GPT-6 Astra More Aligned but Markedly Less Monitorable, Raising AI Safety Alarms(7 posts)→

Original post →

More from Models

Models channel →