Anthropic models now outpace DeepSeek-Flash in decode speed, but oddly long API TTFT raises server-side rewriting suspicion
teortaxesTex · x · 2026-10-10
Based on benchmark measurements, teortaxesTex notes that depending on how you measure, Anthropic's models are now faster than DeepSeek-Flash — though they show curiously long TTFT on the API.
He suspects this isn't genuine speed: Anthropic may be doing server-side screening/rewriting first, then streaming tokens with an artificial delay, which could also help hide batching. The author caveats it's partly shitposting — the speed figures may be an artifact of how the Cherry client accounts usage.
Overall it's a technical observation plus speculation about frontier-model inference speed and API streaming behavior.
Related event: Anthropic Models Beat DeepSeek-Flash in Speed, Yet Show Odd API Latency(2 posts)→
More from Infra
- System 1 ANE: millisecond AI decisions on Apple's Neural Engine with no text generation — pcuenq · 2026-10-10
- oMLX 0.7.1.dev1 adds decision models and batched prefill, +31% Qwen3.6 decode speed on M5 Max — pcuenq · 2026-10-10
- Bittensor GPU rental network sees 45% spend growth, 106% more rentals in monthly report — markjeffrey · 2026-10-10
- Microsoft doubles down on local AI with Nvidia RTX Spark Surface Ultra priced up to $5,899 — MooseEfficient2151 · 2026-10-10
- Building a pit crew for Grok Bot: frontier model plans, free models grind — alexcovo_eth · 2026-10-10
- AI boom turns into a debt boom: Oracle 5y CDS near record 261bps, implying 20.4% default odds — cyb3rops · 2026-10-10