GPT-6 Astra leaks in benchmark: 279/280 enterprise tasks done, zero hallucination
ryanshrout · x · 2026-09-06
Baxate quotes Signal65's preliminary test results for the rumored GPT-6 Astra (unverified, not officially confirmed):
- At maximum reasoning effort, the model completed 279 of 280 real multi-step enterprise jobs end to end and fabricated nothing on unanswerable retrieval questions.
- Its weighted error count is roughly half that of the previous leaders, Claude Fable 5.1 and Muse Spark 1.3.
- At list API prices (input, cached input, and output), it is also cheaper per correct task than the model it displaces.
The thread's core point: don't evaluate models by $/M tokens alone — "cost per task" or "intelligence per token" is what matters for enterprise deployments.
Related event: Leaked Benchmarks Claim GPT-6 Astra Aces Enterprise Tasks(2 posts)→
More from Models
- GPT6 Astra reasoning tiers tested: Max takes 20 min vs 5 min on medium — bdsqlsz · 2026-09-06
- What comes after GPT-6 Astra: from chatbot to a coworker you leave running for a week — johnseach · 2026-09-06
- Claim: OpenAI killed Sora and Atlas to go all-in on Astra and coding models — sahilypatel · 2026-09-06
- Pangram reveals full technical details: training, dataset, architecture and evals — zainhas · 2026-09-06
- GPT Astra 6 powers an intubation training sim no other model could build — BraydonDymm · 2026-09-06
- "Astra is good but too token-heavy": user calls for a GPT-6-grade Sol — CtrlAltDwayne · 2026-09-06