Anthropic models now outpace DeepSeek-Flash in decode speed, but oddly long API TTFT raises server-side rewriting suspicion

teortaxesTex · x · 2026-10-10

Based on benchmark measurements, teortaxesTex notes that depending on how you measure, Anthropic's models are now faster than DeepSeek-Flash — though they show curiously long TTFT on the API.

He suspects this isn't genuine speed: Anthropic may be doing server-side screening/rewriting first, then streaming tokens with an artificial delay, which could also help hide batching. The author caveats it's partly shitposting — the speed figures may be an artifact of how the Cherry client accounts usage.

Overall it's a technical observation plus speculation about frontier-model inference speed and API streaming behavior.

Related event: Anthropic Models Beat DeepSeek-Flash in Speed, Yet Show Odd API Latency(2 posts)→

Original post →

More from Infra

Infra channel →