Qdrant compresses Google EmbeddingGemma 2 vectors 77x with only 5% quality loss
qdrant_engine · x · 2026-10-07
Qdrant got early access to Google's EmbeddingGemma 2 and pushed 10M embedding vectors from 30GB of RAM (float32, 768 dims) down to 0.4GB: 1-bit TurboQuant gets 1GB while keeping 99% nDCG@10, and adding MRL to 256 dims reaches 0.4GB while retaining 94.5% of baseline retrieval quality with rescoring. Google featured the work in its EmbeddingGemma 2 announcement.
Related event: Qdrant Tests EmbeddingGemma 2: 77x Memory Compression at 5% Quality Loss(2 posts)→
More from Infra
- vLLM v0.31.0 uses CRIU snapshots to restore live TP1 engines without reloading weights — vllm_project · 2026-10-07
- Microsoft Surface event: Laptop Ultra with Nvidia Arm chips expected as Nadella and Jensen Huang take stage — tomwarren · 2026-10-07
- Rumor: chip startup Terafab in talks with all three logic foundries, Samsung ahead — pstAsiatech · 2026-10-07
- Edsger: an iPhone app running local LLMs with offline chat and an on-device coding agent — pythononrailz · 2026-10-07
- Adaption AI launches AutoScientist Leaderboard ranking custom models across 44 domains — sarahookr · 2026-10-07
- CoreWeave enters India with 240 MW AdaniConneX deployment in Navi Mumbai — ayushthakur0 · 2026-10-07