Rabdos Launches Math AI Benchmark; Claude Opus 5 Takes the Lead

AI4Code · x · 2026-08-07

Rabdos introduced the Rabdos Math Index, a holistic benchmark designed to evaluate mathematical reasoning across proofs, solutions, and visual perception.

The benchmark features formalized graduate-level theorems, research-level numeric problems, and questions requiring visual reasoning over figures. In the initial rankings, Claude Opus 5 leads with a score of 46, followed by GPT-5.6 Sol and Claude Fable 5 at 39. No other tested model scored above 25.

Related event: Rabdos Launches Math AI Benchmark, Claude Opus 5 Takes Lead(2 posts)→

Original post →

More from Models

Models channel →