New HSS Benchmark Shows Top AI Models Fail Basic Intuitive Visual Reasoning Humans Find Easy
dustinvtran · x · 2026-10-08
Elorian, in partnership with Scale AI, launched Humanity's Sixth Sense (HSS), a 522-question benchmark targeting implicit visual reasoning — affordance, retrodiction, mechanistic causality, and social norms — that humans handle effortlessly.
GPT 6 Astra, Gemini 3.8 Flash, and Claude Opus 5.5 all failed the sample questions on all three attempts:
- Whether two more books fit on a shelf with visible gaps — all said no
- Which side of the frame a traveler entered from — all wrong
- Whether taut chains bear a plough's drag weight — all claimed they hung slack
- Whether people near an aircraft were crew or visitors — all misjudged
- Whether an engine cage spins with the shaft — all said yes
The takeaway: frontier models can pass the bar exam yet systematically fail intuitive physical and commonsense visual reasoning.
More from Models
- Bengio disputes 'just a sandbox bug' framing of AI agent hacks in FT op-ed — AlexTensor · 2026-10-08
- OpenAI model proves Hilbert's Tenth Problem false over Q, sidestepping 80-year approach — aran_nayebi · 2026-10-08
- What Anthropic's $200 tier changes about choosing between Opus and Sonnet — thursdai_pod · 2026-10-08
- AI flip: it may plan your Boston trip before solving the Riemann hypothesis — jxmnop · 2026-10-08
- OpenRouter's Usage and Spend Charts Have Split: Used Models Aren't the Paid Ones — Spiritual_Skirt_9312 · 2026-10-08
- Your Job Is to Push the Model Slightly Out of Distribution — _Stocko_ · 2026-10-08