OpenAI Researcher: Chain-of-Thought Monitoring Still Works but "Fragile and Trending Negative"
zetalyrae · x · 2026-09-04
Responding to calls for frontier-lab transparency, OpenAI researcher merettm says the computation-graph depth of current frontier models including Astra is within 2x of GPT-4, so chain-of-thought monitoring — which OpenAI has preserved since its first reasoning models — still works and offers a view into how alignment generalizes. He admits the technique is "fragile and unfortunately trending in a negative direction" for non-architectural reasons he'll detail soon, and calls strengthening it a core research goal.
More from Models
- New paper: supervising just 1% of tokens can match full on-policy distillation, 0.1% sometimes suffices — jiank_uiuc · 2026-09-23
- Sparse distillation paper: supervising just 0.1%-1% of tokens can match or beat full OPD — jiank_uiuc · 2026-09-23
- Distillation's real impact on Chinese labs debated: no hard evidence, says Lambert, maybe 1-2 month edge — xeophon · 2026-09-23
- Instinct hit by user-data mixing reports; Muse CEO trolls with a safety promise — alexandr_wang · 2026-09-23
- Computer-use faceoff: Grok skips using the computer and just generates the flower — socialwithaayan · 2026-09-23
- BridgeBench: Grok 4.7 is 50% pricier and 60% slower than Grok 4.6 with no quality gain — socialwithaayan · 2026-09-23