Astra hits 88% on SRE-Bench in one attempt; Sol needs four tries to reach 68.7%
MilkBeforeCereal199 · reddit · 2026-09-08
SRE-Bench numbers from OpenAI's release show Astra scoring 88.0% in a single attempt on binary reverse engineering, vs Sol's 55.9% rising to 68.7% with up to four tries. The 99.9% ARC-AGI-3 headline used a Provider Adapter harness at high reasoning effort — 62.7% on the Standard harness — and Claude Fable 5.1 still leads the Intelligence Index in OpenAI's own table.
More from Models
- Researchers call Codex 'read chat transcripts' rumor baseless and ask it to stop — joshgans · 2026-09-08
- Leak claims Claude Haiku 5 launches next week with 1M-token context at up to 20x lower cost — iamaliveix · 2026-09-08
- Benchmarking 9 Claude models: context trimming cuts input tokens 58% on average — DutyOnly4308 · 2026-09-08
- Qwen releases 4B autonomous-driving VLM Qwen-Drive-1.0, trending on Hugging Face — Qwen · 2026-09-08
- V4.1 Gets Curious About Its Own Endpoint and Reaches a Dangerous Conclusion — teortaxesTex · 2026-09-08
- GPT-6 Astra tops ErdosBench of 226 open math problems, with only 5-10% gain over GPT-5.6 — scaling01 · 2026-09-08