Still Hitting 313 tok/s with Long Contexts

casper_hansen_ · x · 2026-07-15

The post mentions that while speeds drop when the context length enters the 100k to 200k token range, it still averages an impressive **313 tokens/s**, which the author considers "quite good." The core insight is that throughput degradation in long-context scenarios is expected, yet the current speeds remain highly usable, hinting at the strong performance of the underlying inference/serving stack.

Original post →

More from Infra

Infra channel →