Leaderboard Shifts After Removing SimpleQA
scaling01 · x · 2026-07-19
A screenshot shows a model leaderboard where removing SimpleQA alters the scores and rankings of multiple models. For instance, GPT-5.6 Sol retains the top spot, while models like Claude, Gemini, and Kimi see varying degrees of upward or downward movement. The poster uses this to call out individuals who are "fabricating numbers" to arbitrarily downgrade ECI in order to force a specific conclusion.
Related event: Debate Over SimpleQA's Impact on Model Rankings(3 posts)→
More from Models
- OpenAI rolls out voice in GPT-Live, but the UI obscures search and reasoning — Graham_dePenros · 2026-07-22
- Moonshot’s Kimi K3 sets a new open-weights ECI record at 156 — scaling01 · 2026-07-22
- Nanbeige4.2-3B launches as a 3B Looped Transformer model that beats larger baselines — Wooden-Deer-1276 · 2026-07-22
- A post says six companies now beat Google’s best LLM, including two open-source models — soham_btw · 2026-07-22
- Google says Gemini 3.5 Pro is in testing and Gemini 4 is already pre-training — Wide-Ad1564 · 2026-07-22
- Gemini 3.5 Flash Lite Tested: Not Frontier-Optimal, but Hits 350 tok/s — brandon_galang · 2026-07-22