Daily probes across 34 LLM APIs catch DeepSeek reasoner's silent 10x token spike

Electrical_Rip892 · reddit · 2026-09-30

Since Aug 20 the author has run the same private probe set daily against 34 models from 15 labs, comparing each model only with its own past — no rankings, no LLM judges, code-graded checks hashed into a public transparency log. First catch: on Sept 10, DeepSeek's reasoner started using roughly 10–12x the thinking tokens it had for the previous three weeks on the same probes, with no announcement — a silent cost/latency change users would only feel as "off." The other 33 models show no quiet capability drops so far; Opus 5.5 has been stable since launch. Code: piperoll/seismograph; live readings: seismo.piperoll.org.

Original post →

More from Models

Models channel →