GPT-6 Astra says you can delete scaffolding — just not the control kind

Slight_Republic_4242 · reddit · 2026-09-15

OpenAI shipped GPT-6 Astra with a bold pitch: proactive agent workflows of 5+ hours saw a 20% pass-rate improvement with fewer inference calls, and the company says its improved decision-making lets you "remove scaffolding." Capability numbers are strong — Terminal-Bench 4.0 at 57.9% (up from 37.3%) and computer use fast enough to finish Financial Modeling World Cup challenges 4x faster than the human winner.

But the safety section is all relative rates: "unintended outcomes" are down 89% vs GPT-5.6 Sol and 74.7% vs Claude Fable 5.1 — with no absolute rate per thousand runs, the only number that matters when the failure mode is deleting your data.

The contradictions stack up: Astra is the first model to hit the Critical cybersecurity threshold under OpenAI's own Preparedness Framework, and enterprise access ships off by default, with site/app allowlists, confirmation gates for consequential actions, and automated review of unauthorized tool calls. The author's key distinction: capability scaffolding (retries, step checks) is a workaround for a weak model and can go; control scaffolding (authorization boundaries, audit trails) is a workaround for a model that can act — and it becomes more necessary as models improve.

Original post →

More from coding & agent

coding & agent channel →