Ling 3.1 Flash posts 33% Terminal-Bench and 62% AutomationBench, huge agentic gains
ArtificialAnlys · x · 2026-10-06
Artificial Analysis published agentic benchmark results for Ling 3.1 Flash: 33% on Terminal-Bench v4.0 (up from 0% for Ling 3.0 Flash) and 62% on AutomationBench-AA (up from 3%), marking large agentic improvements.
Related event: Ant Group's Ling 3.1 Flash Doubles Intelligence Index to 41(4 posts)→
More from Models
- CoNLL 2023 paper: instruction-tuned GPT models beat children on Theory of Mind tests — dioscuri · 2026-10-06
- Mistral now testable in Battle and Agent Modes on LMArena — arena · 2026-10-06
- User Claims 'GPT-6' Solved His Favorite CTF Fully Autonomously in About an Hour — SIGKITTEN · 2026-10-06
- Mistral Large 4 burns over 2x the output tokens per task vs GPT-6 sol — haider1 · 2026-10-06
- Mistral insider hails Large 4 release as finally making product plans come together — qtnx_ · 2026-10-06
- Reflection AI Claims Beam Is 3-4x More Inference-Efficient Than GLM 5.2 — ChrSzegedy · 2026-10-06