GPT-6 Astra ships with API defaulting to low reasoning despite 'Highest reasoning' marketing
docdavkitty · reddit · 2026-09-08
OpenAI's new flagship GPT-6 Astra (1.05M context, $10/$50 per M tokens) looks fine on benchmarks, but the implementation details matter more:
- API defaults to low reasoning: reasoning.effort now has five levels (low/medium/high/xhigh/max) and the API defaults to low, while the model card markets "Highest reasoning" — integrations that don't explicitly raise effort silently buy the cheapest thinking mode.
- Pricing traps: cached input is $1 (0.1×) but cache writes cost $12.50 — a 1.25× surcharge over uncached input; prompts over 272K tokens trigger a cliff billing the whole request at 2× input/cache and 1.5× output.
- Benchmarks by job: DeepSWE 74.1 is a small hop over GPT-5.6 Sol's 72.7; gains concentrate in computer use & agents (BrowseComp 91.5, OSWorld 2.0 72.6). OpenAI highlights Terminal-Bench-Science 64.6 vs Claude Fable 5.1's 52.6.
- Cyber is gated: ExploitBench hits 100, but Astra sits at the Critical cybersecurity threshold under the Preparedness Framework, with advanced defender workflows routed via Trusted Access / Daybreak Blue.
- Wavy rollout: the Sep 3 "limited customers first" appearance looked like an accidental leak, with the formal launch 24h later.
The open question: at 2.5× Sol's price, does "finishes the work" hold only if you manually turn the effort dial up?
More from Infra
- NVIDIA's new Sol-H3 fast inference method for H3 awaits a ComfyUI port — krigeta1 · 2026-09-08
- Walking the AI rack optical stack: InP substrates and silicon photonics as the cleaner bet — demian_ai · 2026-09-08
- Running dual RX 7900 XTX on X570/X870 Taichi for local LLM inference: is x8/x8 enough? — espece-de-bon · 2026-09-08
- Hugging Face teases WebGPU inference engine with 5-10x speedups on Transformers.js — nicodotdev · 2026-09-08
- Meta to deep-dive recommendation inference systems at PyTorch Conference 2026 — PyTorch · 2026-09-08
- Memory crunch hits home: 4TB portable SSD prices stun as AI reprices the storage stack — demian_ai · 2026-09-08