MindTopo benchmark: best MLLM scores 54.1% vs 97.4% human on topological reasoning; GPT-6 Astra needs 8+ hours per task

yining_hong · x · 2026-09-16

Original post →

More from Models

Models channel →