Private visual reasoning bench eyebench saturated by Astra at 97%, author rules out data contamination
adonis_singh · x · 2026-10-09
Developer adonissingh's private visual reasoning benchmark eyebench (spot-the-difference, maze following, topology) once stumped top vision models but has been saturated by Astra. To rule out training-on-API-questions, he generated a fresh set of similar questions; Astra scored marginally higher at 97%, indicating genuine task understanding. He also notes Haiku scores identically on xhigh and max, making fable-5.1's price premium look even starker.
More from Models
- Ramp data shows ElevenLabs leading Voice AI model adoption among businesses — lukeharries · 2026-10-09
- Are URM and Universal Transformers the forgotten architecture beating standard LLMs? — moschles · 2026-10-09
- Grok Bot can now search and monitor X posts to track breaking news and trends — Polymarket · 2026-10-09
- User: Sol 6.1 and Astra became 'really dumb', claim tasks done but don't do them — Yamapama · 2026-10-09
- Commentator: Claude-style dashboards will kill a quarter of data-aggregation SaaS — michalmalewicz · 2026-10-09
- Claude religiously follows robots.txt, and this dev finds it hilarious — generativist · 2026-10-09