GLM-5.3 Scores 28.8% on New RealSWE Benchmark, Closing In on GPT-6 Astra

zainhas · x · 2026-09-12

Specific Labs' new Real-SWE benchmark, which evaluates frontier models on private real-world enterprise codebases, shows Fable 5.1 at 38.8%, GPT-6 Astra at 33.8%, and GLM-5.3 at 28.8%. The author notes GLM-5.3 is surprisingly within spitting distance of the frontier closed models, suggesting open-weight coding models are narrowing the gap.

Related event: Real-SWE Benchmark Tests Frontier Models on Private Enterprise Code(2 posts)→

Original post →

More from Models

Models channel →