Thread questions which model providers can actually sustain 150+ tokens/sec
DanielLockyer · x · 2026-08-04
Users are comparing model/provider combinations for sustained throughput, with the original point being that many services can briefly hit over 150 tokens/sec, then quietly fall back to around 50 tokens/sec the following week.
The thread frames fast models as the fun models and asks what options actually exist for consistently high token generation speed, highlighting a recurring gap between demo performance and real-world serving limits.
More from Infra
- Agentic AI Triggers a Storage Shock: Enterprise Data Becomes the New Bottleneck — BenBajarin · 2026-08-04
- RunPod Test: Generating 30s Video with MiniMax H3 on A100 Costs Just $0.72 — shadowtheimpure · 2026-08-04
- RTX 5090 Inference Test: Capping Power at 480W Costs Less Than 3% Performance — WonderfulEagle7096 · 2026-08-04
- DeepSeek V4 Hits 1,328 tok/s Prefill on Single RTX 6000 via Krasis — mrstoatey · 2026-08-04
- Mach-1 Additive: 35B Model Runs at 120 t/s on Laptops Using 1.7-bit Weights — pbaylies · 2026-08-04
- Surgery on open-weights models: Optimizing inference with hand-rolled Rust implementations — doodlestein · 2026-08-04