OpenBMB's MiniCPM5-2B tops sub-4B open models with Intelligence Index score of 15
ArtificialAnlys · x · 2026-09-07
Artificial Analysis reports that OpenBMB's MiniCPM5-2B scores 15 on the Intelligence Index v4.2—the highest of any open-weights model under 4B total parameters.
Key findings
- A 2.6B dense reasoning model, text-only, 131k context, Apache 2.0.
- Its score of 15 beats Granite 4.2 3B (11) by 4 points, matches Qwen3.5 4B Reasoning (14) with 44% fewer params, and ties Qwen3.5 9B Reasoning (15) at 4x its size.
- Agentic strength: GDPval-AA v2 Elo of 831 leads sub-4B models (110 ahead of Ling 3.0 Tiny at 718); joint-first on τ³-Banking (21%); second on AA-Briefcase (438).
- Weak spots: knowledge, coding, and long context—9% on Humanity's Last Exam, 9% on Terminal-Bench v2.1, 0% on CritPt.
- Low hallucination: attempts only 29% of AA-Omniscience questions, yielding a 78% Non-Hallucination Rate and a -12 score instead of heavy penalties.
- Token-efficient: just 19k output tokens per task (11k reasoning), lowest in the comparison set—relevant for on-device and edge deployment.
Related event: OpenBMB Releases MiniCPM5-2B, Tops Sub-4B Open Model Intelligence Index(4 posts)→
More from Models
- Thread ranks frontier coding AI: Opus 4.5 at average SWE, next-gen above experts — menhguin · 2026-09-07
- GPT-6 Astra Pro tops Simple-Bench at 86.5%, above the human baseline — koltregaskes · 2026-09-07
- Fable 5.1 takes 2nd place on MathArena, but GPT Astra stays better and cheaper — gabrielchua · 2026-09-07
- User finds exact eBay pump listing with Gemini while Claude said it didn't exist — HankYeomans · 2026-09-07
- Reddit Skeptic: Astra Looks Like Catch-Up, Not a Leap — 'AGI Moment' Smells Like Hype — miltonian3 · 2026-09-07
- OpenRouter token usage tracks 27x y/y growth, becoming a key AI demand signal — jeff_weinstein · 2026-09-07