Private benchmarks: Xiaomi v2.6 Pro rated worst model in two years, new Grok only Luna-level
Afinetheorem · x · 2026-09-22
Researcher Afinetheorem argues private benchmarks matter because public ones are far easier to hillclimb than to show genuine general intelligence.
On his own benchmark, the new Xiaomi v2.6 Pro is literally the worst model he has tested in two years, and the new Grok performs at Luna level — far behind top models.
More from Models
- Mimo V2.6 undercuts Grok 4.7 by 6x on output price amid same-day model launches — op7418 · 2026-09-22
- Math community weighs in on AI 'Bel' claims: 100 solved problems, Millennium Problem skepticism — avaitopiper · 2026-09-22
- Terminal-Bench 4.0 leaderboard refresh draws attention to who's on top — ns123abc · 2026-09-22
- How Tencent Hunyuan packed a 770B model into 214 GiB with 5-bit-per-4-weights quantization — TencentHunyuan · 2026-09-22
- xAI fixes SDK bug dropping reasoning content, significantly boosting Grok 4.7 — ns123abc · 2026-09-22
- TTS leaderboard: xAI hits 87.6% pronunciation accuracy, Kokoro 82M fastest at 242 chars/s — ArtificialAnlys · 2026-09-22