Weaviate 1.39's 4-bit Rotational Quantization cuts 10M OpenAI embeddings from 61GB to 7.8GB RAM

philipvollet · x · 2026-09-08

Weaviate 1.39's 4-bit Rotational Quantization shrinks 10M OpenAI embeddings from 61GB to 7.8GB of RAM (6144 → 784 bytes per 1536-dim vector, an honest 7.84x since the 16-byte header is constant). Fast Walsh-Hadamard rotation spreads information evenly; per-vector min/max scalar quantization stores 2 dims per byte; a new asymmetric scheme stores 4-bit but queries at 8-bit, recovering most accuracy at zero storage cost.

Original post →

More from Infra

Infra channel →