OpenAI researcher: GPT-6's CoT controllability keeps rising over RL training
gleech · x · 2026-09-04
OpenAI researcher Tomasz Korbak shared an ongoing finding: for GPT-6, chain-of-thought controllability has been increasing over the course of RL training — which wasn't the case for previous models — and it correlates strongly with no-CoT capabilities across several model generations.
His team is root-causing the increase. Sharer 1a3orn calls it "a pretty big and not great update," underscoring mounting concerns about how reliably CoT can be monitored as models scale.
More from Safety
- Why AI Watermarking May Break Down in Agentic Workflows — yaakg25 · 2026-09-04
- Cheap model writes 700 solid words; jailbreak "tax" drops from $50 to near zero — ctjlewis · 2026-09-04
- Will the EU AI Act Extend to Humanoid Robot Regulation? — DueFoxTheFifth · 2026-09-04
- Forethought weighs a superintelligent "nightwatchman" aboard galactic colonization probes — willmacaskill · 2026-09-04
- Analysis of the "Hugging Face Attack" Extrapolates Rogue AI Agent Scenarios — OK_The_Nomad · 2026-09-04
- GPT-6 Astra reportedly scores 100% on ExploitBench, finds two zero-days in testing — VraserX · 2026-09-04