LLM calorie benchmark: only 16-48% of meals estimated within 20% error

mr_tolkien · reddit · 2026-09-07

The author benchmarked LLMs on estimating calories from 25 Nutrition5k meal photos plus descriptions, scoring models by the share of meals within 20% error. Results: Muse Spark 1.3 led (48% within 20%, median 45 kcal), DeepSeek v4 Flash Vision hit 40%, Qwen 3.8 Flash 36%, while Qwen 3.8 27b managed only 16% (median 148 kcal), badly beaten by the similar-size Muse Glimmer 30b (32%). Key takeaway: rankings don't track model size — the "best" 32GB-VRAM local model is highly task-dependent.

Original post →

More from Models

Models channel →