The Marshmallow AI Benchmark: Testing vision models with counting

qu1etus · reddit · 2026-08-21

The author created the "Marshmallow Benchmark": a photo of marshmallows on a baking sheet, asking AI to count them accurately without guessing. Despite strict prompts, major models (Gemini, Claude, GPT) gave varied results ranging from 472 to 539. It's a fun test highlighting vision model capabilities regarding detail counting and hallucinations.

Original post →

More from Multimodal

Multimodal channel →