Qdrant 实测 EmbeddingGemma 2:1000 万向量内存从 30GB 压到 0.4GB
qdrant_engine · x · 2026-10-07
Google DeepMind 发布首个原生多模态端侧 embedding 开源模型 EmbeddingGemma 2,统一文本、代码、图像、音频与视频到共享向量空间。Qdrant 团队获早期访问,实测其 Matryoshka 兼容 embedding 的压缩极限。
- 1000 万文档的 float32 全尺寸向量需 30.7GB 内存
- 1-bit TurboQuant 量化后约 1GB,保留 99% nDCG@10
- 叠加 MRL 降至 256 维 + rescoring,仅需约 0.4GB,保留 94.5% 基线检索质量
即以约 77 倍的内存节省换取约 5% 的质量损失。该实测被 Google 收进 EmbeddingGemma 2 官方公告,Qdrant 发布了完整技术深潜文章。
所属事件:Qdrant 实测 EmbeddingGemma 2:内存压缩 77 倍(2 条相关)→
「Infra」频道最新
- 英伟达估 NVLink Fusion 市场规模至 2030 年前接近 2500 亿美元 — Beth_Kindig · 2026-10-07
- 博主预言成真:OpenAI 推 500 美元档,企业疯狂控 token 成本 — labeveryday · 2026-10-07
- vLLM XPU 支持上线,Intel Arc Pro B70 本地跑 LLM 有了新引擎 — vllm_project · 2026-10-07
- 晒出自制微型集群:M5 Max 128GB+Linux RTX 6000 跑本地 AI — pcuenq · 2026-10-07
- RTX 6000+M5 笔记本异构跑 MiMo 2.6 Flash,40 tokens/sec — pcuenq · 2026-10-07
- RTX 6000+M5 笔记本异构跑 MiMo 2.6 Flash,40 tokens/sec — pcuenq · 2026-10-07