Vision+Logic Benchmark: Gemini Leads, GPT Stagnates
Afinetheorem · x · 2026-08-13
On a custom vision and logic benchmark, the author found Grok 4.6 and Qwen3.8 Max to be mediocre, scoring between Gemini 2.0 Flash Lite and GPT 5.6 Luna. Gemini models remain the best at this type of task, while GPT models, despite being decent, haven't significantly improved since the o3 release over a year ago.
More from Models
- Grok 4.6 Model Card Analysis: Big Internal Gains, Lags on Public SWE Evals — scaling01 · 2026-08-13
- Anthropic's Leaked Deck Predicts 2026 as the Last Window to Catch Up in AI — imjustnewatai · 2026-08-13
- Researcher Criticizes Frontier Models for Hiding Reasoning Traces, Calls for Open AI Science — rao2z · 2026-08-13
- AI Safety Researcher Points Out Typos and Rushed Third-Party References in Day-1 Model Cards — Miles_Brundage · 2026-08-13
- SemiAnalysis Slams NVIDIA: Committee-Based Frontier Model Development Does Not Work — teortaxesTex · 2026-08-13
- Elon Musk Hints at Grok 4.6, Claiming 'Pareto Gold' Dominance — shaunmmaguire · 2026-08-13