Your p99 latency benchmark may be lying: a deep dive into coordinated omission
Franc0Fernand0 · x · 2026-09-11
The author's second latency article explains why the most common benchmarking approach—waiting for each response before sending the next request—inadvertently coordinates with the system under test and hides real tail latency, a problem known as coordinated omission.
- Example: a 100ms GC pause would queue dozens of slow requests in a real workload, but the benchmark logs just one and moves on, so p99 looks great while lying
- Covers sources of latency variance from CPU caches to garbage collectors, and how latency compounds across services
- Includes a small ping tool to measure your own network's tail latency
More from Infra
- llama.cpp lands Flash Attention tuning for RDNA4, big prefill gains on AMD — pmttyji · 2026-09-11
- Running MiniMax H3 on 12GB VRAM: quantization, Turbo LoRAs and attention backends compared — Possible_Mood676 · 2026-09-11
- Spomin: live KV cache compaction squeezes 500k tokens of context into 180k resident — wgaca2 · 2026-09-11
- PiPNN nearest-neighbor search wins three awards, up to 78x faster index building — khademinori · 2026-09-11
- M.2-Oculink eGPU Link Silently Downgrades to PCIe Gen1 — Here's How to Check — El_90 · 2026-09-11
- DeepSeek launches V4.1-Flash with 1M-token context and 4x smaller KV-cache — matlabulous · 2026-09-11