Scale AI's HSS benchmark: top model scores 53.6% on intuitive visual reasoning vs 93.1% for humans

geoffwolfe · x · 2026-10-10

Scale AI and Elorian released Humanity's Sixth Sense (HSS), an open-source benchmark of 522 human-crafted image/video tasks probing intuitive visual reasoning — implicit temporal, spatial, social, and abstract structure — with rubric-based judging. The best model (GPT-6-astra, max reasoning) hits 53.6% vs 93.1% for humans; median is 30.9%. Models average 4,000 reasoning tokens per task (overthinking without correctness), video tasks drop 7.3 points for 23/25 models, and social understanding is the weakest domain (24.4% avg). Dataset and eval harness are on Hugging Face.

Related event: Scale AI's HSS benchmark: humans score 93.1%, best AI model only 53.6% on intuitive visual reasoning(8 posts)→

Original post →

More from Multimodal

Multimodal channel →