LMArena analyzed 30,086 answer pairs: different LLMs share just 43.1% of ideas
arena · x · 2026-09-12
LMArena's Battle Mode team analyzed 30,086 pairs of real-world Text Arena answers (May 1 – Sep 2) to measure how much conceptual ground models share:
- Different LLMs share only 43.1% of their ideas on average
- Adjacent generations of the same family echo each other: Grok 4.5/4.6 at 59.7%, GPT 5.5/5.6 at 59.2% overlap
- US and Chinese models draw on the same ideas: Fable 5 and GLM 5.3 share 55.9%, with GLM, Kimi, Opus and Sonnet forming a high-overlap neighborhood
Conclusion: the model most similar to Fable is neither Opus nor Sonnet.
More from Models
- Startup reportedly builds autonomous drone system using GPT-6 Astra to track people from a single image — Polymarket · 2026-09-12
- Nex-N2.5 Pro, a 397B multimodal model focused on Computer Use, quietly lands on OpenRouter — nikola_mr64990 · 2026-09-12
- GPT-6 fails to improve on molecular property prediction, fueling AGI skepticism — GaryMarcus · 2026-09-12
- User complains Grok Bot keeps getting dumber with use — billyjhowell · 2026-09-12
- Every Builds Internal Platform for Personal Benchmarks Based on Real Work, Not Leaderboards — danshipper · 2026-09-12
- Gemini live search repeatedly failing on basic queries, users report — Dimensional-Misfit · 2026-09-12