GLM-5.3 Scores 28.8% on New RealSWE Benchmark, Closing In on GPT-6 Astra
zainhas · x · 2026-09-12
Specific Labs' new Real-SWE benchmark, which evaluates frontier models on private real-world enterprise codebases, shows Fable 5.1 at 38.8%, GPT-6 Astra at 33.8%, and GLM-5.3 at 28.8%. The author notes GLM-5.3 is surprisingly within spitting distance of the frontier closed models, suggesting open-weight coding models are narrowing the gap.
Related event: Real-SWE Benchmark Tests Frontier Models on Private Enterprise Code(2 posts)→
More from Models
- Grok 4.7 Was Promised This Week, Now xAI Says It 'Needs More Time to Cook' — alejandroll10 · 2026-09-12
- Grok 4.7 was promised this week, now xAI says it needs more time — draginol · 2026-09-12
- Lost in the middle confirmed in practice: context ordering matters as much as inclusion — ClickOk5811 · 2026-09-12
- GPT-6 Astra review: stunning at 3D games and computer use, still not a daily driver — petergyang · 2026-09-12
- Researcher Challenges 'Verifiability' Intuition: RLVR Generalizes via Manufactured Information Asymmetry — kalomaze · 2026-09-12
- First quantitative evidence: Claude and GPT now use GUIs as well as APIs — ysu_nlp · 2026-09-12