Mistral ML4 Benchmarks: SOTA on Finance and Legal Workflows, Terminal and Spreadsheet Navigation
GuillaumeLample · x · 2026-10-06
Part 4 of the ML4 thread: matches the best open-weight models on DeepSWE, AutomationBench, and AA-Briefcase, and is SOTA on finance/legal workflows and complex multimodal grounding benchmarks. It navigates complex terminal workflows, works across spreadsheets, slides, and PDFs, and reasons over scientific and multimodal tasks.
Related event: Mistral Unveils Mistral Large 4, a 1T-Parameter Open Multimodal Model(42 posts)→
More from Models
- Google rolls out Nano Banana 2.1 with 4K output at $0.076 per image — testingcatalog · 2026-10-06
- Space Bunny Alpha lingers on OpenCode a few days longer than OpenRouter — sull · 2026-10-06
- Analysis: US Models Are Now Cheaper per Task Than Chinese Ones, Upending a Widely Held Assumption — sebkrier · 2026-10-06
- GPT-6.1 Sol spotted in Copilot 365 Premium, ChatGPT rollout likely next — koltregaskes · 2026-10-06
- Are models only improving at verifiable domains? AI stories winning prizes spark debate — erikphoel · 2026-10-06
- Reflection AI's upcoming open-weight model could bring the open-source crown back to the US — The AI Daily Brief · 2026-10-06