Still Hitting 313 tok/s with Long Contexts
casper_hansen_ · x · 2026-07-15
The post mentions that while speeds drop when the context length enters the 100k to 200k token range, it still averages an impressive **313 tokens/s**, which the author considers "quite good." The core insight is that throughput degradation in long-context scenarios is expected, yet the current speeds remain highly usable, hinting at the strong performance of the underlying inference/serving stack.
More from Infra
- A detailed GPU upgrade guide compares H200 and B200 with benchmark code — StasBekman · 2026-07-21
- NVFP4 speeds up Flux, Qwen-Image and other media models in ComfyUI tests — Certain-Will-2769 · 2026-07-21
- Huawei’s Ascend 950DT could beat Nvidia B300 on tokens per watt, analysis says — teortaxesTex · 2026-07-21
- OpenRouter says dynamic routing saved 22,000 users over $100,000 on GLM 5.2 — gajesh · 2026-07-21
- Chinese LLM vendors push API prices lower as competition intensifies — sen_o · 2026-07-21
- Gated Delta Networks paper links Mamba-style memory updates to stronger trillion-scale models — vista8 · 2026-07-21