DoPR: reusable compressed document prefixes cut LLM reranking latency up to 8x

_reachsumit · x · 2026-09-04

DoPR decouples offline document processing from online LLM reranking: query-independent document representations are converted into compressed prefix states precomputed offline and reused across queries, so online scoring only processes the query and scoring token.

On TREC DL, BEIR and BRIGHT with Qwen3 models (0.6B–8B), it achieves up to 8.0x online document-side memory reduction and 8.04x latency speedup while retaining 97.1%–99.5% of full-document rerankers' average NDCG@10.

Original post →

More from Research

Research channel →