Weaviate 1.39's 4-bit Rotational Quantization cuts 10M OpenAI embeddings from 61GB to 7.8GB RAM
philipvollet · x · 2026-09-08
Weaviate 1.39's 4-bit Rotational Quantization shrinks 10M OpenAI embeddings from 61GB to 7.8GB of RAM (6144 → 784 bytes per 1536-dim vector, an honest 7.84x since the 16-byte header is constant). Fast Walsh-Hadamard rotation spreads information evenly; per-vector min/max scalar quantization stores 2 dims per byte; a new asymmetric scheme stores 4-bit but queries at 8-bit, recovering most accuracy at zero storage cost.
More from Infra
- Intel extends High NA EUV lead as TSMC and Samsung confirm adoption by end of decade — BenBajarin · 2026-09-08
- User seeks a custom GGUF quant to fit GLM on a 192 GB RAM Mac between Q2 and Q4 — CentrifugalMalaise · 2026-09-08
- Community GPU guide compares GB per dollar and memory bandwidth across cards — jacek2023 · 2026-09-08
- Report claims 1 million high-NA EUV wafers per year, industry insider says it's possible — pstAsiatech · 2026-09-08
- Qwen3-0.6B (400MB) on a 2017 Galaxy Note 8 Drives Real Desktop Chrome via Structured Page Perception — Mean-Standard7390 · 2026-09-08
- SEMICON Taiwan 2026: AI Chips Shift From Node Scaling to System Scaling With Six Battlegrounds — pstAsiatech · 2026-09-08