SWE-Bench Pro V2 launches already saturated, sparking benchmark fatigue jokes
burny_tech · x · 2026-09-23
ScaleAI Labs announced SWE-Bench Pro V2 with a thread of updates — and X users immediately joked that the new benchmark launches already saturated, underscoring how frontier models outpace software-engineering benchmark design.
Related event: SWE-Bench Pro V2 Hard Hits 98% on Day One, Instantly Saturated(2 posts)→
More from Models
- Marco Polo Benchmark Shows Chinese Model Retention Falling, Sparking Questions — teortaxesTex · 2026-09-23
- Yuchen Jin: Opus 5.5 underwhelms, frontier LLM coding has plateaued — Yuchenj_UW · 2026-09-23
- Forward Future puts Opus 5.5 through 8 tests: cities, games, animation — MatthewBerman · 2026-09-23
- Matthew Berman: Opus 5.5 is the best model in the world — MatthewBerman · 2026-09-23
- Alibaba's Ovis-Embedding tops five benchmarks with unified omni-modal embeddings — _reachsumit · 2026-09-23
- A $5, 10-minute SFT run boosts Qwen3.6 by 8-12% on GPQA and MMLU-Pro — simonguozirui · 2026-09-23