Chinese model's 74.2 score under fire: best of 8 eval variants, maxed thinking budget, ~2.5x cost

teortaxesTex · x · 2026-09-10

A detailed critique of a new Chinese model's reported benchmark results highlights three issues: it ran the same exam under 8 different evaluation setups and reported the best one (74.2 vs an average of 69.0, range 65.5–74.2) while competitors ran once — its claimed 74.2 vs Opus 5's 74.0 win only holds under someone else's harness, dropping to 72.6 with its own default. It also maxed the thinking tier (+8 points but 2.5x cost, admitted by its own report as inefficient), so the headline score isn't reproducible at the recommended everyday tier, undercutting the "cheap" pitch. Training also used these benchmarks as the progress metric with cherry-picked checkpoints. The author concedes it isn't cheating — all vendors report best scores and the full 8-way table was published — and verified the architecture's efficiency claims (890 bytes/token, 8B active) independently check out.

Original post →

More from Models

Models channel →