PUMA Sparsifies Multimodal Embeddings, Cuts Storage 16x, Speeds Up Retrieval 25x

_reachsumit · x · 2026-08-27

PUMA introduces a post-hoc sparsification recipe using a TopK autoencoder to convert frozen multimodal embeddings into compact sparse codes without retraining the backbone. The method preserves dense dot-product geometry during pretraining before fine-tuning the sparse encoder for retrieval. Evaluations on Qwen3-VL-Embedding-2B show that PUMA matches or outperforms dense retrieval on four out of five benchmarks, reducing vector storage by 8-16x and achieving up to 25x faster exact dense scoring on larger candidate pools.

Original post →

More from Infra

Infra channel →