Chimera Boosts Multi-Vector Retrieval Throughput by 16x via GPU-CPU Co-Processing

_reachsumit · x · 2026-08-25

Chimera is a new multi-vector retrieval system designed to address bottlenecks in existing approaches through GPU-CPU co-processing. Unlike prior GPU-based systems like PLAID, which are limited by data transfer overhead during queries, Chimera stores highly compressed, low-precision quantization codes on the GPU while maintaining high-precision data in CPU memory.

By leveraging GPU-resident data for efficient candidate generation and filtering, and employing a collaborative scoring scheme, Chimera completely avoids vector data transfer and enables computation overlap. Experiments on real-world datasets demonstrate that Chimera achieves up to 16.0x higher Queries Per Second (QPS) than existing methods at the same recall level.

Original post →

More from Infra

Infra channel →