Top AI Models Score Under 50% on New Math Figure Reasoning Benchmark

prof_g · x · 2026-08-06

Rabdos AI introduced Vizi Bench, a benchmark evaluating mathematical figure reasoning with 27 problems.

Results reveal significant shortcomings in current leading AI models: the best score was only 48.1%, while the bottom model scored just 3.7%. This indicates that most models have yet to learn how to 'see' and reason over diagrams like a mathematician.

Original post →

More from Models

Models channel →