Scale AI's HSS benchmark: top model scores 53.6% on intuitive visual reasoning vs 93.1% for humans
geoffwolfe · x · 2026-10-10
Scale AI and Elorian released Humanity's Sixth Sense (HSS), an open-source benchmark of 522 human-crafted image/video tasks probing intuitive visual reasoning — implicit temporal, spatial, social, and abstract structure — with rubric-based judging. The best model (GPT-6-astra, max reasoning) hits 53.6% vs 93.1% for humans; median is 30.9%. Models average 4,000 reasoning tokens per task (overthinking without correctness), video tasks drop 7.3 points for 23/25 models, and social understanding is the weakest domain (24.4% avg). Dataset and eval harness are on Hugging Face.
More from Multimodal
- Alibaba's TaoMate-H3 streams minute-long video with audio using 3-step denoising — jiqizhixin · 2026-10-10
- SGLang-Diffusion serving framework for diffusion models to be unveiled at PyTorch Conference 2026 — PyTorch · 2026-10-10
- image-blaster repo turns a single photo into a fully explorable 3D world in five minutes — anthara_ai · 2026-10-10
- MoMA commissions Holly Herndon and Mat Dryhurst for dual-venue AI art show — matdryhurst · 2026-10-10
- Reference Videos Make Minimax Acting Believable: Full Prompt Template Shared — roychodraws · 2026-10-10
- Open-source creative harness Voyager adds 8 languages, drives Blender, Resolve and 100+ tools — ravisparikh · 2026-10-10