8 Vision Models Tested on 2,000 Animal-Trace Photos; Best Species Accuracy Just 37.15%
d_kielbasa · reddit · 2026-09-29
A Reddit user built Wildlife CSI, benchmarking 8 vision models on 2,000 animal-trace photos — 400 each of droppings, footprints, feathers, eggs, and bones. The best model reached only 37.15% species accuracy.
- Footprints were the weakest category for every model: Opus got 205/400 eggs right but only 82/400 footprints.
- Across all eight runs, 842 photos received no correct species guess from any model.
- Caveats: no expert-human baseline; blocked/failed requests (Azure blocked some GPT images) stay in the denominator; species are unevenly represented and reasoning effort doesn't equalize compute.
Photos come from the AnimalClue datasets by Risa Shinoda et al., labels checked against iNaturalist observations. Write-up and code are open-sourced.
More from Models
- One env var change broke ticket tagging: GLM 5.3 no longer accepts thinking off — sweetcake_1530 · 2026-09-29
- H Company launches Holo4, a family of generalist computer-use models for desktop, web and Android — RemiCadene · 2026-09-29
- Bindu Reddy calls Muse Spark a benchmaxxed Claude distillation, says DeepSeek Flash is way better — bindureddy · 2026-09-29
- Anthropic and OpenAI ship five model releases in one month as 'pace the frontier' pact collapses — XFreeze · 2026-09-29
- OpenAI internal models show ~15-min 80% time-horizon on real research tasks, far below METR — ben_j_todd · 2026-09-29
- One Claude session generated a full video explaining Attention, burning 271K context tokens — Abhishekcur · 2026-09-29