The Marshmallow AI Benchmark: Testing vision models with counting
qu1etus · reddit · 2026-08-21
The author created the "Marshmallow Benchmark": a photo of marshmallows on a baking sheet, asking AI to count them accurately without guessing. Despite strict prompts, major models (Gemini, Claude, GPT) gave varied results ranging from 472 to 539. It's a fun test highlighting vision model capabilities regarding detail counting and hallucinations.
More from Multimodal
- Seedance 2.5 vs Kling 3.0 Omni: A Comparison of Expressiveness and Character Interaction — SimplyAnnisa · 2026-08-21
- LTX-2.3-10Eros_I2V Trends on Hugging Face — amisima · 2026-08-21
- Tencent WithEveryone Enables Identity-Preserving Group Gen — Tencent-Hunyuan · 2026-08-21
- How to Prevent Background Music in Minimax H3 Generated Videos? — Early-Reputation3186 · 2026-08-21
- Minimax H3 Creates Stunning Video, Users Amazed — Eric520CC · 2026-08-21
- Depth Anything V4 Uses Riemannian Flow Matching for Dynamic 4D Scene Reconstruction — RexDouglass · 2026-08-21