LLM calorie benchmark: only 16-48% of meals estimated within 20% error
mr_tolkien · reddit · 2026-09-07
The author benchmarked LLMs on estimating calories from 25 Nutrition5k meal photos plus descriptions, scoring models by the share of meals within 20% error. Results: Muse Spark 1.3 led (48% within 20%, median 45 kcal), DeepSeek v4 Flash Vision hit 40%, Qwen 3.8 Flash 36%, while Qwen 3.8 27b managed only 16% (median 148 kcal), badly beaten by the similar-size Muse Glimmer 30b (32%). Key takeaway: rankings don't track model size — the "best" 32GB-VRAM local model is highly task-dependent.
More from Models
- GPT-6 Astra fails badly at dialog-centric interpersonal game's first-time user experience — msew · 2026-09-07
- GPT 6 Astra Strains Under Heavy Usage as Users Pushed to Switch to 5.6 sol — vista8 · 2026-09-07
- Rumor: Gemini 3.5 Live may be coming soon, possibly Pro-subscription only — Able-Line2683 · 2026-09-07
- ACX: Claude Fable 5.1 Cracks 370-Year-Old Cipher, Apollo Research Expands Hiring — Astral Codex Ten · 2026-09-07
- OpenAI's Astra Decompiles BIOS to Nail EC Backlight Control Other Models Couldn't — anom604 · 2026-09-07
- User reports Astra high feels better and burns far less weekly quota than Sol — sweetbeard · 2026-09-07