OpenAI serving only 30 TPS? Dev says local Qwen at 50 TPS now matches official speed
drdanielbender · x · 2026-10-10
- Developer Daniel Bender is surprised OpenAI serves only 30 tokens/sec, quoting thsottiaux that reaching 50 TPS with the most optimized tokenizer makes efficient models even faster in practice.
- His local experiments with Qwen3.8-flash-next hit 50 TPS, putting local inference on par with OpenAI's official serving speed.
More from Infra
- Bittensor miners squeeze up to 54% more inference throughput from the same GPUs — bittingthembits · 2026-10-11
- SkyPilot meetup recap: GPU scheduling pain, Holo4 training on K8s, and SGLang's roadmap — skypilot_org · 2026-10-11
- PyTorchCon talk: MoRI + vLLM brings RDMA KV-cache transfer and wide expert parallelism to AMD — PyTorch · 2026-10-11
- ODS V3 pre-release turns any PC, Mac or Linux box into a private AI server — tom_doerr · 2026-10-10
- Drex 1.5 lands on OpenRouter, billed as the fastest decision model available — Div_pradeep · 2026-10-10
- Opening a Site to AI Agents: 820 Crawler Hits, 280 MCP Connections, 3 Real Tool Calls — S_B_B_B_K · 2026-10-10