Qdrant squeezes EmbeddingGemma 2 vectors to 0.4GB from 30GB, keeping 94.5% quality
qdrant_engine · x · 2026-10-07
Google DeepMind released EmbeddingGemma 2, its first natively multimodal on-device embedding open model unifying text, code, images, audio and video. Qdrant got early access and tested compression limits: 10M float32 vectors (30.7GB RAM) drop to 1GB with 1-bit TurboQuant (99% nDCG@10 kept), and to 0.4GB adding MRL to 256 dims with rescoring (94.5% quality kept) — a 77x memory saving for 5% quality loss. Google featured the work in its official announcement; full deep dive on Qdrant's blog.
Related event: Qdrant Tests EmbeddingGemma 2: 77x Memory Compression at 5% Quality Loss(2 posts)→
More from Infra
- vLLM v0.31.0 uses CRIU snapshots to restore live TP1 engines without reloading weights — vllm_project · 2026-10-07
- Microsoft Surface event: Laptop Ultra with Nvidia Arm chips expected as Nadella and Jensen Huang take stage — tomwarren · 2026-10-07
- Rumor: chip startup Terafab in talks with all three logic foundries, Samsung ahead — pstAsiatech · 2026-10-07
- Edsger: an iPhone app running local LLMs with offline chat and an on-device coding agent — pythononrailz · 2026-10-07
- Adaption AI launches AutoScientist Leaderboard ranking custom models across 44 domains — sarahookr · 2026-10-07
- CoreWeave enters India with 240 MW AdaniConneX deployment in Navi Mumbai — ayushthakur0 · 2026-10-07