Vision+Logic Benchmark: Gemini Leads, GPT Stagnates

Afinetheorem · x · 2026-08-13

On a custom vision and logic benchmark, the author found Grok 4.6 and Qwen3.8 Max to be mediocre, scoring between Gemini 2.0 Flash Lite and GPT 5.6 Luna. Gemini models remain the best at this type of task, while GPT models, despite being decent, haven't significantly improved since the o3 release over a year ago.

Original post →

More from Models

Models channel →