Private visual reasoning bench eyebench saturated by Astra at 97%, author rules out data contamination

adonis_singh · x · 2026-10-09

Developer adonissingh's private visual reasoning benchmark eyebench (spot-the-difference, maze following, topology) once stumped top vision models but has been saturated by Astra. To rule out training-on-API-questions, he generated a fresh set of similar questions; Astra scored marginally higher at 97%, indicating genuine task understanding. He also notes Haiku scores identically on xhigh and max, making fable-5.1's price premium look even starker.

Original post →

More from Models

Models channel →