New HSS Benchmark Shows Top AI Models Fail Basic Intuitive Visual Reasoning Humans Find Easy

dustinvtran · x · 2026-10-08

Elorian, in partnership with Scale AI, launched Humanity's Sixth Sense (HSS), a 522-question benchmark targeting implicit visual reasoning — affordance, retrodiction, mechanistic causality, and social norms — that humans handle effortlessly.

GPT 6 Astra, Gemini 3.8 Flash, and Claude Opus 5.5 all failed the sample questions on all three attempts:

The takeaway: frontier models can pass the bar exam yet systematically fail intuitive physical and commonsense visual reasoning.

Related event: Scale AI's HSS Benchmark: Top Models Lag Far Behind Humans on Intuitive Visual Reasoning(2 posts)→

Original post →

More from Models

Models channel →