Gemini Leads in Vision and Logic Tasks, Nearing 50% Pass@1 Accuracy

Afinetheorem · x · 2026-08-14

Despite mixed public reception, Gemini remains the clear leader in vision and logic tasks, especially in video understanding, offering top-tier performance at its price point. The author notes that models are getting ever closer to a 50% pass@1 rate on these evaluations, with some achieving exact matches on 73% of cases.

Related event: Gemini Leads in Visual-Logic Evaluation as GPT Hits Bottleneck(2 posts)→

Original post →

More from Models

Models channel →