GPT-6 System Card's log-scale graph obscures Astra's CoT controllability jump, Reddit user argues
thegamebegins25 · reddit · 2026-09-04
Reddit user thegamebegins25 argues a chart in OpenAI's GPT-6 Astra System Card is "bordering on deceptive":
- The log-scale y-axis visually downplays a major change in chain-of-thought controllability
- The underlying data shows Astra maintains near-100% CoT controllability up to a few hundred tokens of CoT length — i.e., the model can control its own chain of thought to defeat monitors
- Models from earlier in the summer could only control their CoT 10% of the time
This matters because CoT monitoring is currently one of our strongest safety monitors; if models can control their CoT far more often, monitor effectiveness drops sharply. Source: OpenAI's deploymentsafety page, figure 28.
More from Models
- Unverified leak claims OpenAI ARC-AGI-3 score jumped from 7.8% to 98.6% — miilesus · 2026-09-04
- GPT-6 ties with Grok 4.6 and Muse Spark 1.3 on benchmark — ns123abc · 2026-09-04
- OpenAI's GPT-6 Astra hits Critical cybersecurity threshold, first model to do so — moyix · 2026-09-04
- GPT-6 Astra crushes ARC-AGI-3: 62.7% standard, 99.9% with adapter harness — Hesamation · 2026-09-04
- Astra Beats GPT Pro at Proof-Checking, a First for Non-Pro Models — joshgans · 2026-09-04
- tszzl: Astra's capabilities remain far from fully explored — tszzl · 2026-09-04