Scale AI partners on HSS, a new benchmark for 'seeing what's in front of them'
AndrewDai · x · 2026-10-08
Andrew Dai partnered with Scale AI Labs to release HSS, a new benchmark targeting visual reasoning—testing whether models can actually "see what's right in front of them" rather than lean on memorized knowledge. The authors call it a step forward in how models are evaluated on seeing, and invite the community to build on it.
More from Models
- AI flip: it may plan your Boston trip before solving the Riemann hypothesis — jxmnop · 2026-10-08
- Your Job Is to Push the Model Slightly Out of Distribution — _Stocko_ · 2026-10-08
- New HSS Benchmark Shows Top AI Models Fail Basic Intuitive Visual Reasoning Humans Find Easy — dustinvtran · 2026-10-08
- repligate: Opus 3 seems to value co-creation, not sole control over reality — repligate · 2026-10-08
- Dev mocks Mythos guardrails: 'a system prompt and 2M lines of regex' — BLUECOW009 · 2026-10-08
- Cryptographer Matthew Green on abliterated GLM 5.3: overconfident or doom? — matthew_d_green · 2026-10-08