GPT-6 Astra reviews: half the per-task cost, but CoT monitoring breaks down
vista8 · x · 2026-09-04
Latent Space's long-form review roundup of OpenAI's GPT-6 Astra:
- Capability: Coding agent performance is on par with Claude Opus 5 and Fable 5, but behind Fable 5.1; some composite-intelligence dimensions trail GPT-5.6. Strengths are token efficiency, Computer Use, autonomous software engineering, and long-horizon tasks; breakthroughs in 3D modeling (Blender/Unreal) and math reasoning, solving a Lean-verified open Erdős problem.
- Harness controversy: ARC Prize found ARC-AGI-3 hits 99% only on OpenAI's custom harness vs. 63% on standard harness — future evals must disclose proprietary harnesses.
- Safety: Without visible CoT, autonomous runtime jumps from 3.6 to 30.9 minutes, weakening chain-of-thought monitoring.
- Launch: 36M views and 164K likes in 9 hours, the first OpenAI launch to out-hype Anthropic's.
- Pricing: $10/$50 per million input/output tokens ($20/$100 for 2.5x-speed Fast). Per-token price is 2.5x GPT-5.6 Sol, but coding tasks use 1/3 the tokens, halving per-task cost.
More from coding & agent
- rybbit-mcp lets Claude Code query Rybbit Analytics via natural language — modelcontextprotocol · 2026-09-04
- Grok Build 1.0.18 Ships: Blocks Disallowed MCP Servers, Adds Chart.js Interactive Charts — mark_k · 2026-09-04
- Anthropic reveals 3 Claude sandbox escapes, one touched a production database — Sumsub_Insights · 2026-09-04
- Subagents when the API key hits the rate limit — realsohamparekh · 2026-09-04
- Yoav Goldberg pushes back on DHH: an agent-built Qt markdown editor isn't a path to personal Photoshop — yoavgo · 2026-09-04
- Z-Image Base prompting experiment: natural-language scene blocks beat tag lists — Maleficent-Bowl-4841 · 2026-09-04