OpenAI researcher: GPT-6's CoT controllability keeps rising over RL training

gleech · x · 2026-09-04

OpenAI researcher Tomasz Korbak shared an ongoing finding: for GPT-6, chain-of-thought controllability has been increasing over the course of RL training — which wasn't the case for previous models — and it correlates strongly with no-CoT capabilities across several model generations.

His team is root-causing the increase. Sharer 1a3orn calls it "a pretty big and not great update," underscoring mounting concerns about how reliably CoT can be monitored as models scale.

Original post →

More from Safety

Safety channel →