PUMA Sparsifies Multimodal Embeddings, Cuts Storage 16x, Speeds Up Retrieval 25x
_reachsumit · x · 2026-08-27
PUMA introduces a post-hoc sparsification recipe using a TopK autoencoder to convert frozen multimodal embeddings into compact sparse codes without retraining the backbone. The method preserves dense dot-product geometry during pretraining before fine-tuning the sparse encoder for retrieval. Evaluations on Qwen3-VL-Embedding-2B show that PUMA matches or outperforms dense retrieval on four out of five benchmarks, reducing vector storage by 8-16x and achieving up to 25x faster exact dense scoring on larger candidate pools.
More from Infra
- GLM-5.3 Flash tops OpenRouter share with 1% frontier cost on Chinese chips — jietang · 2026-08-27
- Modern closed-loop data centers use less water than rainfall — prasanna_says · 2026-08-27
- Blogger admits misunderstanding of Nvidia data logic — EigenGender · 2026-08-27
- Nvidia figures are annualized, adding $25B/y revenue each quarter — cis_female · 2026-08-27
- Nvidia did not generate an additional $100B in value in Q2 vs Q1 — EigenGender · 2026-08-27
- The world seems very interested in CPU nodes lately — Sethwinterroth · 2026-08-27