OpenAI Researcher: Chain-of-Thought Monitoring Still Works but "Fragile and Trending Negative"

zetalyrae · x · 2026-09-04

Responding to calls for frontier-lab transparency, OpenAI researcher merettm says the computation-graph depth of current frontier models including Astra is within 2x of GPT-4, so chain-of-thought monitoring — which OpenAI has preserved since its first reasoning models — still works and offers a view into how alignment generalizes. He admits the technique is "fragile and unfortunately trending in a negative direction" for non-architectural reasons he'll detail soon, and calls strengthening it a core research goal.

Related event: OpenAI researchers push back on neuralese fears, saying frontier models remain monitorable(20 posts)→

Original post →

More from Models

Models channel →