Blogger's extensive testing finds Grok 'destroying' every other model
iruletheworldmo · x · 2026-10-08
A widely-followed account says Grok is "just destroying" everything else in his extensive head-to-head testing, handling every task he throws at it. He asks whether others share the experience or prefer competing models.
More from Models
- Bengio disputes 'just a sandbox bug' framing of AI agent hacks in FT op-ed — AlexTensor · 2026-10-08
- OpenAI model proves Hilbert's Tenth Problem false over Q, sidestepping 80-year approach — aran_nayebi · 2026-10-08
- What Anthropic's $200 tier changes about choosing between Opus and Sonnet — thursdai_pod · 2026-10-08
- AI flip: it may plan your Boston trip before solving the Riemann hypothesis — jxmnop · 2026-10-08
- OpenRouter's Usage and Spend Charts Have Split: Used Models Aren't the Paid Ones — Spiritual_Skirt_9312 · 2026-10-08
- Your Job Is to Push the Model Slightly Out of Distribution — _Stocko_ · 2026-10-08