Thread questions which model providers can actually sustain 150+ tokens/sec

DanielLockyer · x · 2026-08-04

Users are comparing model/provider combinations for sustained throughput, with the original point being that many services can briefly hit over 150 tokens/sec, then quietly fall back to around 50 tokens/sec the following week.

The thread frames fast models as the fun models and asks what options actually exist for consistently high token generation speed, highlighting a recurring gap between demo performance and real-world serving limits.

Original post →

More from Infra

Infra channel →