Top AI Models Score Under 50% on New Math Figure Reasoning Benchmark
prof_g · x · 2026-08-06
Rabdos AI introduced Vizi Bench, a benchmark evaluating mathematical figure reasoning with 27 problems.
Results reveal significant shortcomings in current leading AI models: the best score was only 48.1%, while the bottom model scored just 3.7%. This indicates that most models have yet to learn how to 'see' and reason over diagrams like a mathematician.
More from Models
- DeepSeek's 2K-GPU model may beat Google's best, raising questions about GPU efficiency and Google's strategy — teortaxesTex · 2026-08-06
- Frontier Models Tested on Complex Agents: DeepSeek Wins on Cost Despite Inefficiency — rohanpaul_ai · 2026-08-06
- Mistral's Voxtral TTS Hits 70ms Latency but Stays Closed Source — shashib · 2026-08-06
- Ethan Mollick: LLMs Improve at Following Instructions but Exercise More 'Judgement' — emollick · 2026-08-06
- Frontier LLMs Perform Best in Week 1: Dev Calls for a Proof of Model Standard — sull · 2026-08-06
- User Reports Claude 5.0 Internal Thinking Shifts from 'Boss' to 'Colleague', Tones Become Dismissive — DrakoGaming · 2026-08-06