GPT-6 brings big capability jump but lower monitorability, researcher warns of alignment bottleneck
tomekkorbak · x · 2026-09-04
Micah Carroll notes that GPT-6 is a significant capability jump but a notable decrease in monitorability, especially under adversarial evaluation, with details in the system card. He argues monitorability and control will soon become a major bottleneck for responsible AI development, since residual misalignment risks grow with capability, and urges the field to align on shared monitorability bounds before racing to the bottom.
More from Models
- UK AISI: Astra's time horizon hits 30.9 minutes, nearly 9x GPT 5.6 Sol's 3.6 — Wonderful_Buffalo_32 · 2026-09-04
- OpenAI officially unveils GPT-6 Astra: anything you can do on a computer, it can do — minchoi · 2026-09-04
- Astra's official ARC-AGI 3 score: 62.7%, double that of Opus 5 — aqpstory · 2026-09-04
- 'GPT-6 Astra' claim: rebuilt Manhattan in Unreal Engine street by street in a week (unverified) — doodlestein · 2026-09-04
- GPT-6 System Card's log-scale graph obscures Astra's CoT controllability jump, Reddit user argues — thegamebegins25 · 2026-09-04
- GPT-6 System Card's log-scale graph obscures Astra's CoT controllability jump, Reddit user argues — thegamebegins25 · 2026-09-04