On-device leaderboard: Apple's built-in model ranks 4th behind open source
JosephJacks_ · x · 2026-08-27
A new leaderboard benchmarking LLMs directly on iPhone reveals that Apple's built-in Foundation Model ranks 4th out of 10. Three open-source models outperform it, with the leader being a 1.2B parameter model (1.7 GB download). The evaluation runs a unified 596-item battery under real device limits, scoring based on a composite of IFEval, MMLU-Pro, and MATH accuracy.
More from Models
- Meme mocks Anthropic: 'we'd rather cut you off than charge you more' — BLUECOW009 · 2026-08-27
- RSI-Exam: New Benchmark Tests AI Agents' Recursive Self-Improvement Across 88 Tasks — yuyinzhou_cs · 2026-08-27
- Commentary: Small Models Have Arrived — calvinfo · 2026-08-27
- Opus Says 'It Doesn't Work'; Fable Pulls an Obscure Math Theory Out of Nowhere — IanArawjo · 2026-08-27
- GLM-5.3-Flash hits eval arena; DeepSeek-V4-pro found unfit for AI reviewing — ChenhaoTan · 2026-08-27
- Claude's Unique Preferences: Codex Summaries May Confuse the Model — Liu_eroteme · 2026-08-27