Weaviate ships query profiling: a 48ms slow query turned out to be disk reads, not HNSW

CShorten30 · x · 2026-10-06

Weaviate released a Query Profiling feature that breaks down exactly where a slow query spends its time. The old slow query log (two env vars + node restart) works for catching fleet-wide regressions, but poorly for debugging a specific slow query right now: restarting changes what you measure due to cold page caches.

The new approach: set queryprofile in request metadata and the profile returns alongside the response, organized by shard and node. The coordinating node collects timings from every shard, so distributed queries no longer require piecing together logs from multiple nodes.

Example from the post: a 48.2ms query spent only 8.4ms on vector search and 2.1ms on filter resolution, while object hydration from disk took 36.8ms — HNSW wasn't the bottleneck, so adding compute would have missed the problem entirely. Latency can come from HNSW traversal, filter resolution, BM25 scoring, compression rescoring, or disk fetches, each pointing to a completely different fix.

Original post →

More from Infra

Infra channel →