Ant Group's finance-focused open model Ling-3.0-flash-Fin scores 23 on AA Intelligence Index, trails its VL sibling
ArtificialAnlys · x · 2026-09-16
Artificial Analysis benchmarked Ling-3.0-flash-Fin, Ant Group's new finance-focused open-weights model built on Ling-3.0-flash and co-developed with financial institutions.
Key findings:
- Scores 23 on the Intelligence Index and 24 on the Finance & Accounting Index (same as the VL variant), sitting on the intelligence-vs-active-parameters Pareto frontier with only 5.1B active of 124B total parameters.
- On AA-Briefcase (complex business workflows producing spreadsheets, decks, memos), Fin scores 967 Elo vs VL's 986, passing fewer rubric checks (23.5% vs 24.9%) and lower analytical quality (866 vs 907) — yet higher presentation Elo (1095 vs 1076), surprising given it lacks image input.
- On GDPval-AA v2 (professional knowledge work), Fin scores 1171, 50 points behind VL (1225) but above MiniMax-M2.7 (1087).
- Fin has higher business-knowledge accuracy (17% vs 11%) but worse hallucination (67% vs 81% non-hallucination rate), and outputs 67k tokens per task — 34% more than VL's 50k.
- Both Ling models sit below the total-parameter frontier: Qwen3.8 27B (xhigh) scores 34 with only 27B total parameters.
Related event: Ant's Open-Source Ling-3.0-flash-Fin Benchmarked by Artificial Analysis(3 posts)→
More from Venture
- Arcee AI raises Series B at over $1B valuation to push open-source Trinity models — julsimon · 2026-09-16
- How do you cap an AI agent's API spend across multiple vendors? — apyhubnico · 2026-09-16
- BackOps AI raises $42M Series B six months after $26M Series A to automate logistics ops — HankYeomans · 2026-09-16
- Typeless cuts free tier from 8K to 2K weekly as big models build in voice input — xiaohu · 2026-09-16
- Cohere and Aleph Alpha merge to form first transatlantic foundation model developer with 1,000+ staff — cohere · 2026-09-16
- Dev turns $87k/month business into a game where real revenue grows the city — eptwts · 2026-09-16