OpenAI researcher warns frontier model depth is within 2x of GPT-4 as CoT monitoring erodes
connoraxiotes · x · 2026-09-03
- OpenAI researcher @merettm warns against a race into "unmonitorability," noting the computation graph depth of current frontier models (including Astra) is within a factor of two of GPT-4.
- He says OpenAI has preserved and used chain-of-thought monitoring since its first reasoning models, viewing it as a window into how alignment generalizes from training distribution — but calls the technique fragile and trending negative for reasons unrelated to architecture, with more to come.
- Quoting @tylertracy321: the performance-vs-serial-depth graph makes it tempting to crank depth while creatively ignoring the monitorability-vs-depth graph — the exact incentive structure worth worrying about.
More from Models
- New paper: supervising just 1% of tokens can match full on-policy distillation, 0.1% sometimes suffices — jiank_uiuc · 2026-09-23
- Sparse distillation paper: supervising just 0.1%-1% of tokens can match or beat full OPD — jiank_uiuc · 2026-09-23
- Distillation's real impact on Chinese labs debated: no hard evidence, says Lambert, maybe 1-2 month edge — xeophon · 2026-09-23
- Instinct hit by user-data mixing reports; Muse CEO trolls with a safety promise — alexandr_wang · 2026-09-23
- Computer-use faceoff: Grok skips using the computer and just generates the flower — socialwithaayan · 2026-09-23
- BridgeBench: Grok 4.7 is 50% pricier and 60% slower than Grok 4.6 with no quality gain — socialwithaayan · 2026-09-23