Qdrant releases 10B-vector retrieval dataset with exact top-1000 ground truth for 100K queries
qdrant_engine · x · 2026-09-08
Qdrant has released a 10.07B-vector retrieval benchmark dataset on Hugging Face, providing exact top-1000 ground truth for 100K queries, along with Supernova, an open-source benchmark engine for internet-scale vector datasets.
The authors argue current vector search benchmarks have three structural flaws:
- They top out between 10M and 100M embeddings
- They lack ground truth search results
- They ignore sparse and multi-vector representations used in production hybrid search and filtering
The project aims to push the community beyond micro-optimizations on curated million-scale corpora toward architectures built for billion-vector complexity, building on HF-hosted embedded datasets like Cohere's 250M-vector Wikipedia dump.
More from Infra
- JapanFold launches: 8 open biology AI models served free for Japanese researchers — DavidBennett__ · 2026-09-08
- Running a local MLX music model via Codex pushes M4 MacBook Air to its limits — Dimillian · 2026-09-08
- Debunking the viral 'data center infrasound harm' videos: citations don't hold up — AndyMasley · 2026-09-08
- AI cluster builders bypass the grid: gas turbines, fuel cells and SMRs come onsite — AccBalanced · 2026-09-08
- Jensen Huang calls compute a rentable asset as H100 rental prices jump 22% to $3.28/hr — AccBalanced · 2026-09-08
- Inside the VMs Powering Mobile AI Agents: Instinct, Claude Code — RohanAdwankar · 2026-09-08