Nanbeige and LFM2.5 tie for top mobile model score of 63

ArtificialAnlys · x · 2026-08-25

At the standard 16K context limit, Nanbeige4.2-3B (Reasoning) and LFM2.5-2.6B (Reasoning) tie for the top average score at 63. Individual highlights include Qwen3.5 9B leading in BFCL and GPQA Diamond, Falcon-H1R-7B achieving 97% on MATH-500, and LFM2.5-2.6B showing significantly lower hallucination rates. Nanbeige4.2-3B won no single evaluation but ranked top-five in BFCL, GPQA Diamond, and MATH-500, securing its overall tie.

Related event: Artificial Analysis and Liquid AI Launch On-Device Small Model Benchmarks(8 posts)→

Original post →

More from Models

Models channel →