Astra scores 61 on independent index despite 97.6% FrontierMath hype, at 2.5x the price
eyishazyer · x · 2026-09-04
A close look at the gap between Astra's launch-day numbers and independent evaluations:
- Real wins: 97.6% on FrontierMath Tier 4, 96.0% on GPQA Diamond, 100% on ExploitBench.
- But tied independently: Artificial Analysis' Intelligence Index gives Astra 61 — the same as Sol, the model it replaced, and below Fable 5.1 (66) and Muse Spark 1.3.
- Harness gaming: the headline 99.9% on ARC-AGI-3 came from a $19K custom harness; the default harness yields 62.7% at higher cost.
- Pricing: $10/M input, $50/M output — 2.5x Sol's cost, matching Fable 5.1 despite an equal intelligence score.
- Genuine progress: better token efficiency (valuable for agents) and the first model to cross OpenAI's Critical cybersecurity threshold.
The kicker: Brockman said "welcome to the AGI era" on stage, then downgraded it to a "mission concept" under questioning — hours before the independent numbers landed.
Related event: GPT-6 Astra reviews: strong gains but benchmark claims questioned(3 posts)→
More from Models
- Qwen3.8-Max-0902 coding training lifts RSI-Exam recursive self-improvement score 22% — HuaxiuYaoML · 2026-09-04
- MazeBench 3D environment is now free to play online — patience_cave · 2026-09-04
- GPT-5.6 Sol tops MazeBench; Fable 5.1 could take the lead — patience_cave · 2026-09-04
- Gemini Flash agents improve world modeling: 0% to 4% in two months on MazeBench — patience_cave · 2026-09-04
- MazeBench results: Gemini 3.8 Flash scores 4%, most models under 1% in 3D open world — patience_cave · 2026-09-04
- GPT-6 Astra looks less like AGI and more a specialist in computer-use agents — becomingengageably · 2026-09-04