Qdrant compresses Google EmbeddingGemma 2 vectors 77x with only 5% quality loss

qdrant_engine · x · 2026-10-07

Qdrant got early access to Google's EmbeddingGemma 2 and pushed 10M embedding vectors from 30GB of RAM (float32, 768 dims) down to 0.4GB: 1-bit TurboQuant gets 1GB while keeping 99% nDCG@10, and adding MRL to 256 dims reaches 0.4GB while retaining 94.5% of baseline retrieval quality with rescoring. Google featured the work in its EmbeddingGemma 2 announcement.

Related event: Qdrant Tests EmbeddingGemma 2: 77x Memory Compression at 5% Quality Loss(2 posts)→

Original post →

More from Infra

Infra channel →