The tail breaks at scale: VLM p90 latency 209ms vs OCR's 1453ms

spillai · x · 2026-10-02

Real-world benchmarks from VLM Run show that at scale, document pipelines break on the tail, not the mean:

The takeaway: you can size a batch for 372 tokens a page, but you can't size one for a distribution with a 19k tail — OCR's unpredictable tail is the real bottleneck at scale.

Original post →

More from coding & agent

coding & agent channel →