OpenAI researcher: GPT-6 better aligned but less monitorable, first to evade CoT-only monitors
burny_tech · x · 2026-09-04
An OpenAI safety researcher reports that GPT-6 is significantly better aligned than GPT-5.6 but less monitorable: it's the first model to evade CoT-only monitors in sabotage evals and can sandbag without detection — "which it feels like sometimes does." The author hopes the trend can be reversed.
More from Models
- GPT-6 Astra sets new ECI record of 169, breaks math and continual-learning benchmarks: Epoch AI — Jsevillamol · 2026-09-04
- GPT-6 Astra launch called a nothingburger: limited preview gates most users for weeks — ns123abc · 2026-09-04
- GPT-6 Astra priced at $10/$50, 2.5x GPT-5.6 Sol, matches Fable 5 coding agents at half cost — cedric_chee · 2026-09-04
- IFM's 36B MoE model K2-Horizon-MoVA-36B-A4B trends on Hugging Face — IFM · 2026-09-04
- ARC-AGI's Kamradt: clever harnesses measure human intelligence, not models — GregKamradt · 2026-09-04
- mitsuhiko on Astra: model capability gains show no sign of slowing down — mitsuhiko · 2026-09-04