Probing 34 LLM APIs daily since August caught DeepSeek's silent 3x cost jump on reasoner

Electrical_Rip892 · reddit · 2026-10-01

Since Aug 20, a developer has run identical private probes daily against 34 models from 15 labs over production APIs, grading with plain code (no LLM judge) and timestamping every reading into Rekor.

What it caught: on Sept 10 — the day DeepSeek shipped V4.1 Flash — deepseek-reasoner jumped to 10-12x its usual thinking tokens on two probe sets, with no changelog line about the reasoner alias. Same answers, more tokens, slower and pricier: daily cost went from 7 cents to 25 and has sat around 3x since.

What it hasn't caught: any quiet capability drop across all 34, Opus 5.5 included — read that as "nothing big," not "nothing."

Two reusable tips: split probes into a fixed set plus a regenerated set from the same task families (the gap between them is itself a signal — one model scores 1.00 fixed vs 0.76 fresh), and track thinking tokens as a first-class metric since they move earlier and harder than accuracy. Tooling is MIT-licensed, data CC BY.

Related event: Dev Probes 34 LLM APIs Daily, Catches DeepSeek Silent Changes(2 posts)→

Original post →

More from coding & agent

coding & agent channel →