GPT-6 Astra Pro tops Simple-Bench at 86.5%, above the human baseline
koltregaskes · x · 2026-09-07
GPT-6 Astra and Astra Pro are now on the Simple-Bench leaderboard: Astra Pro scores 86.5%, surpassing the human baseline, while Astra scores 83.6%.
Commenter koltregaskes argues the human baseline for Simple-Bench is set quite high, and models probably passed it a while ago.
More from Models
- OpenAI to cut Cursor's model access on Nov 12 following SpaceX acquisition; Google ships TimesFM-3 — thione · 2026-09-07
- Google releases TimesFM-3: 330M-parameter zero-shot forecasting with native multivariate support — thione · 2026-09-07
- Google launches WeatherNext 3, its most advanced weather AI, wired into Search and Maps — thione · 2026-09-07
- Runway Unveils Solaris, an Interface World Model Rendering Interactive Apps Frame by Frame — thione · 2026-09-07
- Alibaba Upgrades Qwen3.8-Max-0902: 2.4T Params, 1M Context, at $2/$6 per Million Tokens — thione · 2026-09-07
- Google Releases Gemini 3.8 Flash and Flash Cyber, Upgrading Its Low-Cost Agentic Model — thione · 2026-09-07