Scale AI's HSS Benchmark: Top Models Lag Far Behind Humans on Intuitive Visual Reasoning
Scale AI and ElorianAI released the HSS benchmark of 522 open-ended image, text and video tasks testing intuitive visual reasoning. Humans scored 93.1 while the best model, GPT-6-astra, managed only 53.6.
2026-10-08 ~ 2026-10-08 · 2 related posts
- Scale AI's new visual intuition benchmark: humans 93.1%, best model GPT-6-astra 53.6% — dustinvtran · 2026-10-08
- New HSS Benchmark Shows Top AI Models Fail Basic Intuitive Visual Reasoning Humans Find Easy — dustinvtran · 2026-10-08