RaBitQ lands in Vectorium: up to 30x vector compression, 1.4-2.2x faster than PQ
CShorten30 · x · 2026-09-15
RaBitQ quantization is now integrated into Vectorium: up to 30x compression with no training required, flexible 1/2/4/8 bits per component, and 1.4-2.2x faster than the original implementation. On neural embeddings it beats PQ on every axis — compression, speed, and accuracy — with the gap widening as dimensionality grows.
More from Infra
- vLLM-Omni: serving voice, video, and diffusion models explained by Red Hat AI — vllm_project · 2026-09-15
- Running Z-Image Turbo locally on an RX 6800: full ROCm setup, 43s per image — AstroFieldsGlowing · 2026-09-15
- Sparse GEMM Deserves Attention: Insights From a Dense-GEMM Optimizer — goyal__pramod · 2026-09-15
- Redditor scores unopened DGX Spark for $4k on Craigslist, $1-2k below retail — michaelthatsit · 2026-09-15
- ThunderKittens Runs on NVIDIA Vera Rubin: NVFP4 GEMM Hits 22 PFLOPS — NVIDIAAI · 2026-09-15
- Should this founder drop €8k on a 5090 workstation or stick with cloud GPUs? — abandonedexplorer · 2026-09-15