Weaviate ships query profiling: a 48ms slow query turned out to be disk reads, not HNSW
CShorten30 · x · 2026-10-06
Weaviate released a Query Profiling feature that breaks down exactly where a slow query spends its time. The old slow query log (two env vars + node restart) works for catching fleet-wide regressions, but poorly for debugging a specific slow query right now: restarting changes what you measure due to cold page caches.
The new approach: set queryprofile in request metadata and the profile returns alongside the response, organized by shard and node. The coordinating node collects timings from every shard, so distributed queries no longer require piecing together logs from multiple nodes.
Example from the post: a 48.2ms query spent only 8.4ms on vector search and 2.1ms on filter resolution, while object hydration from disk took 36.8ms — HNSW wasn't the bottleneck, so adding compute would have missed the problem entirely. Latency can come from HNSW traversal, filter resolution, BM25 scoring, compression rescoring, or disk fetches, each pointing to a completely different fix.
More from Infra
- llama.cpp adds DFlash speculative decoding for Qwen3.8-27B, faster than MTP — victormustar · 2026-10-06
- GLM 5.3 full NVFP4 deployable on 4x B200 or H200 with Marlin kernels — TheZachMueller · 2026-10-06
- OpenAI reportedly spent tens of millions in compute to crack Navier-Stokes in days — haider1 · 2026-10-06
- Grid waits hit 7 years for 100MW in Northern Virginia; Crusoe cuts it to 1 year, valued at $30.9B — import_jmr · 2026-10-06
- Veda's sparse-attention ComfyUI node renders 5s video 2.9x faster on a 12GB RTX 5070 — lmoroney · 2026-10-06
- Should Developers Care What Hardware Runs Their Inference API? — ekhyatt · 2026-10-06