Image Counting Test: Both Gemini and GPT Fail at Counting Objects

bytebot · x · 2026-08-18

Tested Gemini 3.7 Flash and GPT-5.6 Sol High on counting items in images (e.g., stacks of colored dumbbells). Both models were initially incorrect. Asking the model to 'verify' eventually leads to the correct answer, but humans remain faster overall.

Original post →

More from Models

Models channel →