Chinese model's 74.2 score under fire: best of 8 eval variants, maxed thinking budget, ~2.5x cost
teortaxesTex · x · 2026-09-10
A detailed critique of a new Chinese model's reported benchmark results highlights three issues: it ran the same exam under 8 different evaluation setups and reported the best one (74.2 vs an average of 69.0, range 65.5–74.2) while competitors ran once — its claimed 74.2 vs Opus 5's 74.0 win only holds under someone else's harness, dropping to 72.6 with its own default. It also maxed the thinking tier (+8 points but 2.5x cost, admitted by its own report as inefficient), so the headline score isn't reproducible at the recommended everyday tier, undercutting the "cheap" pitch. Training also used these benchmarks as the progress metric with cherry-picked checkpoints. The author concedes it isn't cheating — all vendors report best scores and the full 8-way table was published — and verified the architecture's efficiency claims (890 bytes/token, 8B active) independently check out.
More from Models
- IFM's 375B MoE open-weights model K2-Horizon trends on Hugging Face — IFM · 2026-09-10
- Leak claims DeepSeek v4.1 Flash matches or beats GPT-5.6 Sol and Claude Opus 5 — airesearch12 · 2026-09-10
- A visual guide to reasoning_effort on DeepSeek V4.1 Flash — incarnadine72 · 2026-09-10
- 'Gemini 3.8 Flash' demo claims task completion with self-correction in 3 turns — Artistic_Solution117 · 2026-09-10
- "Alien architecture" model design stuns, blogger suggests layering recurrent depth on top — scaling01 · 2026-09-10
- Users Petition OpenAI for $400-$600 Heavy Builder Tier as $200 Plan Runs Dry in 48 Hours — dragonwarrior_1 · 2026-09-10