SWE-rebench adds five-language coding benchmark, GLM-5.2 tops open weights
Fabulous_Pollution10 · reddit · 2026-07-29
SWE-rebench released a major multilingual update for real-world software engineering tasks across five languages: Go, Java, Python, Rust, and TypeScript.
- The new leaderboard adds a multilingual slice beyond the original Python focus.
- Among open-weight models, GLM-5.2 [high] leads with 62.9% Pass@1 and 81.1% Pass@5.
- Other results shown include MiniMax M3, MiMo V2.5 Pro, DeepSeek-V4 Pro, Qwen3.6-27B, Qwen3.6-35B-A3B, and Qwen3.5-35B-A3B.
- The team says the next update in 3–4 weeks will focus heavily on models suitable for local deployment.
- They are asking for suggestions from people actually using local models for software development or coding agents, and the dataset is available for running your own agents.
More from Research
- Coding agents can speed up projects by quietly locking in the wrong decisions — bibryam · 2026-07-29
- ICLR’s 2026 deadline joke lands with a NeurIPS timing twist — delliott · 2026-07-29
- Microsoft removes Mage Flow and points to a more efficient Qwen-VL-class encoder — Dante_77A · 2026-07-29
- Profluentbio is building protein foundation models for AI-designed therapeutics — nathanbenaich · 2026-07-29
- VisualPatchWorld: Code World Models for Efficient Planning — HKBU-KnowComp · 2026-07-29
- A staged diagnosis finds most short-text generation loss comes from the codec — ITMO · 2026-07-29