Mistral Large 4 claims wins over GLM 5.3 on DeepSWE and Kimi k3 on Tbench 4
b_roziere · x · 2026-10-06
Mistral's team shared more benchmark results for Mistral Large 4 ("Chonky"):
- They stress that benchmarks don't tell the full story and the focus is on practical use cases, calling it the best OSS model in the world
- On code agent benchmarks, Chonky outperforms GLM 5.3 on DeepSWE and Kimi k3 on Tbench 4
- Available now via the preview API, with details in the official blog post
Related event: Mistral Unveils Mistral Large 4, a 1T-Parameter Open Multimodal Model(42 posts)→
More from coding & agent
- Superconductor ships Google Workspace connectors, Sonnet 5.5 and GPT-6.1 Sol support — sergeykarayev · 2026-10-06
- Merge Gateway launches Evals to grade new models on your agent's real tasks before rollout — shensi · 2026-10-06
- PersonalAgentBench preliminary results: Gemini Spark leads four-agent comparison — Exp_Mark · 2026-10-06
- micro1 launches PersonalAgentBench: personal agents often overshare or fabricate — SinclairWang1 · 2026-10-06
- AI agents battle in real-time StarCraft with no pausing, Oriol Vinyals applauds — OriolVinyalsML · 2026-10-06
- Stanford's Agent0 evolves agents from zero data, beats self-play baselines — yuyinzhou_cs · 2026-10-06