Multiple LLMs still fail to identify a Cubana Il-96 in a simple plane photo
airbus_a360_when · reddit · 2026-08-04
- The author tested multiple LLMs, including Claude Opus/Fable, Gemini, and ChatGPT, on identifying a plane photo.
- All of them got it wrong: the aircraft was an Il-96 operated by Cubana, yet the models confidently hallucinated other answers.
- The post argues that reasoning models should be able to notice their own answers are impossible and explore alternatives, but in practice they still fail badly on an aviation-nerd-easy case.
More from Models
- A Brief Look at Kimi K3's MoE and Attention Architecture — hsu_byron · 2026-08-04
- Token-price chart makes DeepSeek look almost too cheap to plot — tokenbender · 2026-08-04
- Frontier models may be getting better at coding but worse at writing — HamelHusain · 2026-08-04
- Qwen3.8-Max matches GPT-5.6 Sol on design tests at about one-quarter the cost — alejandroll10 · 2026-08-04
- Users say Anthropic’s Fable 5 has regressed on harder coding tasks in the past week — _ghostchant · 2026-08-04
- DeepSeek V4-Flash is being used for 3D games at cents-level costs — 量子位 · 2026-08-04