Scale AI's HSS Benchmark: Top Models Lag Far Behind Humans on Intuitive Visual Reasoning

Scale AI and ElorianAI released the HSS benchmark of 522 open-ended image, text and video tasks testing intuitive visual reasoning. Humans scored 93.1 while the best model, GPT-6-astra, managed only 53.6.

2026-10-08 ~ 2026-10-08 · 2 related posts