DoPR: reusable compressed document prefixes cut LLM reranking latency up to 8x
_reachsumit · x · 2026-09-04
DoPR decouples offline document processing from online LLM reranking: query-independent document representations are converted into compressed prefix states precomputed offline and reused across queries, so online scoring only processes the query and scoring token.
On TREC DL, BEIR and BRIGHT with Qwen3 models (0.6B–8B), it achieves up to 8.0x online document-side memory reduction and 8.04x latency speedup while retaining 97.1%–99.5% of full-document rerankers' average NDCG@10.
More from Research
- After Chess and Go: Can Any AI Engine Actually Beat Humans at Scrabble? — zuilserip · 2026-09-20
- Textbook author: 99.9% accuracy can mean zero scientific discoveries — bravo_abad · 2026-09-20
- rasbt: Jev's Secret Sauce Is Data, Not the Algorithm — Laya Rival Falls Far Short — RichmanRonald · 2026-09-20
- Sebastian Raschka open-sources an end-to-end 'AI text detector from scratch' project — rasbt · 2026-09-20
- Sebastian Raschka: restricting LLM outputs to an action space is just a classic encoder classifier — rasbt · 2026-09-20
- Inception Labs CEO Stefano Ermon bets on diffusion LLMs over autoregressive decoding — No Priors · 2026-09-20