Thread ranks frontier coding AI: Opus 4.5 at average SWE, next-gen above experts
menhguin · x · 2026-09-07
In a discussion on how capable frontier models have become, @menhguin offers a tiered take: Claude 3.5 Sonnet already clears the "above the median person" bar, Opus 4.5 roughly clears the "average SWE" bar, and rumored next-gen models ("Astra" or "Fable") may reach "likely better than an expert in a sensitive technical operation" level.
More from Models
- OpenAI to cut Cursor's model access on Nov 12 following SpaceX acquisition; Google ships TimesFM-3 — thione · 2026-09-07
- Google releases TimesFM-3: 330M-parameter zero-shot forecasting with native multivariate support — thione · 2026-09-07
- Google launches WeatherNext 3, its most advanced weather AI, wired into Search and Maps — thione · 2026-09-07
- Runway Unveils Solaris, an Interface World Model Rendering Interactive Apps Frame by Frame — thione · 2026-09-07
- Alibaba Upgrades Qwen3.8-Max-0902: 2.4T Params, 1M Context, at $2/$6 per Million Tokens — thione · 2026-09-07
- Google Releases Gemini 3.8 Flash and Flash Cyber, Upgrading Its Low-Cost Agentic Model — thione · 2026-09-07