The tail breaks at scale: VLM p90 latency 209ms vs OCR's 1453ms
spillai · x · 2026-10-02
Real-world benchmarks from VLM Run show that at scale, document pipelines break on the tail, not the mean:
- p90 end-to-end latency: VLM at 209ms (only +9% over its own p50), vs OCR->Text at 1453ms (+168% over p50).
- Token distribution: VLM consumes a fixed 372 tokens per page; OCR->Text has a 2,002-token median but a 19,159-token worst-case page.
The takeaway: you can size a batch for 372 tokens a page, but you can't size one for a distribution with a 19k tail — OCR's unpredictable tail is the real bottleneck at scale.
More from coding & agent
- steipete: btrfs CoW is great for worktrees but terrible for sqlite — steipete · 2026-10-02
- Google Mantis: a skills pack for security review with coding agents — udmrzn · 2026-10-02
- NVIDIA and Cedana serve a portfolio of coding models on one 8x B200 node — josh_wills · 2026-10-02
- LukeW teases Intent Mobile: agent coordination at the level of intent — LukeW · 2026-10-02
- LukeW: Dev tools haven't reached their final form after Terminal UI and chat clients — LukeW · 2026-10-02
- Dev Benchmarks 10+ Coding Agent Workflows: Astra Planning + Sol 6.1 Beats Pricier Combos — kevinkern · 2026-10-02