Nanbeige and LFM2.5 tie for top mobile model score of 63
ArtificialAnlys · x · 2026-08-25
At the standard 16K context limit, Nanbeige4.2-3B (Reasoning) and LFM2.5-2.6B (Reasoning) tie for the top average score at 63. Individual highlights include Qwen3.5 9B leading in BFCL and GPQA Diamond, Falcon-H1R-7B achieving 97% on MATH-500, and LFM2.5-2.6B showing significantly lower hallucination rates. Nanbeige4.2-3B won no single evaluation but ranked top-five in BFCL, GPQA Diamond, and MATH-500, securing its overall tie.
Related event: Artificial Analysis and Liquid AI Launch On-Device Small Model Benchmarks(8 posts)→
More from Models
- Polymarket Forecast on DeepSeek Pro; Bloomberg Warns of Hacker Use — Polymarket · 2026-08-25
- ComfyUI nodes enable grafting Krea2 features into MiniMax H3 — Key-Philosopher-9327 · 2026-08-25
- Planning $100 benchmark for Qwen quantization and KV cache trade-offs — m_mukhtar · 2026-08-25
- Users Frustrated: LLMs Cite Scholarly Consensus to Dismiss Your Own Claims — louisvarge · 2026-08-25
- OpenCode Go Offers $60 of Usage for a $10 Monthly Sub — chrisalbon · 2026-08-25
- GPT-5.6 Sol Max beats Fable 5 Max on DeepSWE at $6.47 vs $21.63 per task — reach_vb · 2026-08-25