Daily probes across 34 LLM APIs catch DeepSeek reasoner's silent 10x token spike
Electrical_Rip892 · reddit · 2026-09-30
Since Aug 20 the author has run the same private probe set daily against 34 models from 15 labs, comparing each model only with its own past — no rankings, no LLM judges, code-graded checks hashed into a public transparency log. First catch: on Sept 10, DeepSeek's reasoner started using roughly 10–12x the thinking tokens it had for the previous three weeks on the same probes, with no announcement — a silent cost/latency change users would only feel as "off." The other 33 models show no quiet capability drops so far; Opus 5.5 has been stable since launch. Code: piperoll/seismograph; live readings: seismo.piperoll.org.
More from Models
- Codex users cry shrinkflation: cached inputs now billable, usage down to a fifth — StewartalsopIII · 2026-09-30
- Gemini Can't Read Secondary Google Calendars or Custom Task Lists, Users Find — Here4Zipline · 2026-09-30
- Leak claims Google's Astra 6.1 arrives in October — iruletheworldmo · 2026-09-30
- Cohere launches Embed 5, a new family of enterprise embedding models — cohere · 2026-09-30
- Z.ai called China's closest answer to Anthropic, with big domestic compute injection still to come — pstAsiatech · 2026-09-30
- GPT-6.1 Sol review: near-flagship feel with barely-moving usage limits — VraserX · 2026-09-30