New 'Humanity's Sixth Sense' benchmark: humans score 93.1%, best model only 53.6%
Charuru · reddit · 2026-10-08
A new benchmark called Humanity's Sixth Sense tests intuitive visual reasoning spanning spatial and causal reasoning and social understanding.
- Humans average 93.1%, while the median model scores just 30.9%
- The strongest model, GPT-6-astra, reaches only 53.6%
- The gap suggests intuitive visual understanding remains far beyond current models
More from Research
- Hugging Face launches Robotic Episodes Viewer for 24k+ LeRobot datasets — mishig25 · 2026-10-09
- Blind humanoid walks, plays soccer and lifts suitcases with joint encoders only — accepted at Humanoids 2026 — Jan_R_Peters · 2026-10-09
- Delete object info from observations and PPO learns to search anyway — TU Darmstadt on its Humanoids 2026 paper — Jan_R_Peters · 2026-10-09
- U-Space finds an interpretable subspace for LLM uncertainty, no training needed — Tobias Braun · 2026-10-09
- CARE certifies VLA inference speedups up to 10.8x with statistical guarantees — UMCP · 2026-10-09
- SOL: a sample-based distributional metric proposed for evaluating text diffusion LMs — NandoDF · 2026-10-09