Image Counting Test: Both Gemini and GPT Fail at Counting Objects
bytebot · x · 2026-08-18
Tested Gemini 3.7 Flash and GPT-5.6 Sol High on counting items in images (e.g., stacks of colored dumbbells). Both models were initially incorrect. Asking the model to 'verify' eventually leads to the correct answer, but humans remain faster overall.
More from Models
- User reports Anthropic appears to have removed the 50% weekly limit for Fable — legit_api · 2026-08-18
- Anthropic's per-token cost runs 4.4x the Vercel average, and devs keep paying — The Decoder · 2026-08-18
- Seeking top-tier LLM with low refusal rate — nobodyreadusernames · 2026-08-18
- Opinion: DeepSeek should shift focus to larger pre-training instead of 4.1 — zephyr_z9 · 2026-08-18
- Observation: Claude Opus 5 dominates the 3D demo scene — techartist_ · 2026-08-18
- Codex Still Disastrous at Interfaces, Corpus May Be Bootstrapped — wavefnx · 2026-08-18