Scale AI's new visual intuition benchmark: humans 93.1%, best model GPT-6-astra 53.6%
dustinvtran · x · 2026-10-08
Scale AI partnered with ElorianAI to launch "Humanity's Sixth Sense," a benchmark of 522 open-ended image and video tasks testing everyday intuitive visual reasoning — spatial, causal, and social understanding. The gap is stark: humans score 93.1%, the strongest model GPT-6-astra reaches only 53.6%, and the median model scores just 30.9%.
More from Research
- CoLM 2026 Poster: Vibe-Voting LLMs and Why Users Distrust Benchmarks — boknilev · 2026-10-08
- Baseten's Base Labs has all 3 papers accepted at NeurIPS workshops — baseten · 2026-10-08
- Open-vocabulary text search fused into Google 3D Tiles across 27M voxels of San Francisco — bilawalsidhu · 2026-10-08
- SEAR research project to debut at COLM 2026, paper and code coming soon — WenhuChen · 2026-10-08
- LeJEPA lets you pretrain self-supervised vision models on your own data — randall_balestr · 2026-10-08
- LEDGER builds persistent 3D object memory from egocentric video, answers questions without rewatching — mangahomanga · 2026-10-08