SWE-bench Pro leaderboard puts Ornith-1.0-397B and GLM-5.2 at the top
victormustar · x · 2026-07-22
SWE-bench Pro leaderboard update
The screenshot shows the SWE-bench Pro leaderboard on Hugging Face.
- deepreinforce-ai/Ornith-1.0-397B leads with 62.2.
- zai-org/GLM-5.2 is close behind at 62.1.
- poolside/Laguna-S-2.1 ranks third at 59.4.
- Several other models cluster around 59–58 points, including MiniMax-M3, Kimi-K2.6, GLM-5.1, Hy3, and MiMo-V2.5-Pro.
The image is a compact snapshot of how current models are stacking up on a software-engineering benchmark.
More from Models
- Encode Bench finds Base64 output quality tracks AI intelligence scores at 0.91 — Valuable-Repeat-7347 · 2026-07-22
- Usage caps are forcing model downgrades, and users can feel the floor drop — dioscuri · 2026-07-22
- GPT-5.6 health answers reportedly beat human doctors in blinded tests — PeterDiamandis · 2026-07-22
- Kimi founder’s decade-long track record and Moonshot’s real moat — SumitGup · 2026-07-22
- User says DeepSeek v4 flash is good enough to cancel Claude Code — Simple-Status5933 · 2026-07-22
- Claude Fable 5 lands in the middle on a 3D HTML dashboard cost test — Tarandjpop · 2026-07-22