MindTopo benchmark: best MLLM scores 54.1% vs 97.4% human on topological reasoning; GPT-6 Astra needs 8+ hours per task
yining_hong · x · 2026-09-16
- Northwestern, Microsoft Research and Stanford release MindTopo, a topological reasoning benchmark for MLLMs built on 5 Piagetian primitives (continuity, separation, order, enclosure, knots), 13 task types, and 11,016 instances, evaluating 11 MLLMs.
- Key takeaway: frontier models can recognize topological relations in static scenes but cannot maintain or operate on them across action sequences — best model reaches 54.1% vs 97.4% human accuracy.
- The quoted demo shows GPT-6 Astra solving tricky spatial constraint tasks (interlocked parts, threading a rope through three rings), but search took 1h40m for the first task and 8+ hours for the second.
- Author Manling Li argues reasoning is the core bottleneck and highlights whether models can abstract reusable topological structures as transferable skills.
More from Models
- LLaDA-Image: 6B fully-diffusion DiT trained on 90% image-only data, no caption bottleneck — jiqizhixin · 2026-09-16
- GPT-5.6 Sol flips its conclusions when you just ask 'Are you sure?' — Sockand2 · 2026-09-16
- Student finds Opus and ChatGPT trash his slides, then reverse after seeing the source material — Kazoru4 · 2026-09-16
- Altman says internal post-Astra OpenAI model can solve problems the world's best mathematicians cannot — rohanpaul_ai · 2026-09-16
- Agent randomly references a money quote the user never said while editing unrelated text — Kyrannio · 2026-09-16
- OpenAI's new quirk: chat mode now has a usage cap while work mode doesn't — koltregaskes · 2026-09-16