Scale AI's new visual intuition benchmark: humans 93.1%, best model GPT-6-astra 53.6%

dustinvtran · x · 2026-10-08

Scale AI partnered with ElorianAI to launch "Humanity's Sixth Sense," a benchmark of 522 open-ended image and video tasks testing everyday intuitive visual reasoning — spatial, causal, and social understanding. The gap is stark: humans score 93.1%, the strongest model GPT-6-astra reaches only 53.6%, and the median model scores just 30.9%.

Original post →

More from Research

Research channel →