GPT-6 Astra reportedly tops FrontierSWE at 65.5%, nearly doubling predecessor
charliermarsh · x · 2026-09-18
ProximalHQ claims GPT-6 Astra tops the FrontierSWE benchmark with 65.5%, beating Fable 5.1 (56.3%) and its predecessor GPT-5.6 Sol (32.2%). The claim is unverified by OpenAI and should be treated with caution.
Related event: Report: GPT-6 Astra Tops FrontierSWE Benchmark(2 posts)→
More from Models
- RL agents invent their own diagnostic renderings to ground code understanding, sparking RL scaling optimism — teortaxesTex · 2026-09-18
- Dev discovers Codex security hardening switched persistent agent sessions to per-message instances — RileyRalmuto · 2026-09-18
- Astra for Law posts big legal benchmark gains as Mollick asks if labs will eat every AI vertical — emollick · 2026-09-18
- Anthropic's stealth model accused of hardcoded routing to Opus 5 — teortaxesTex · 2026-09-18
- GPT-6 Astra beats Factorio: Space Age in just 2 days — ResultBackground2450 · 2026-09-18
- Self-described ChatGPT co-inventor launches Jev, claiming 20-200x speed at 40-400x lower cost — multiply_matrix · 2026-09-18