JPMorgan paper: pooled LLM eval cuts retrieval selection cost by up to 4.9x
_reachsumit · x · 2026-09-03
JPMorganChase researchers present an incremental pooled LLM evaluation method for cost-effective retrieval model selection in production RAG systems.
- An LLM judges the union of documents retrieved by all candidate systems; as new systems join, only their novel documents are judged, and judgments are reused across systems.
- Validated on 4 retrieval benchmarks with 11 systems (dense/sparse/hybrid) and deployed to compare 62 retrieval configurations for a financial news QA system.
- Rankings correlate strongly with gold-standard evaluation; 97% of pairwise orderings survive bootstrap uncertainty in the qrels.
- In production, document overlap yields 65–80% judgment reuse and up to 4.9x lower evaluation cost, making pooled LLM eval a practical workflow for incremental retrieval selection.
More from Research
- HCI papers increasingly use LLM judges while obfuscating it, researcher warns — IanArawjo · 2026-09-03
- Why VRChat Particle Pools Still Work: Gravity Naturally Converges States — Michael_Moroz_ · 2026-09-03
- Radix Sort Hit 5B Key-Value Pairs per Second With Zero Compute Shaders — Michael_Moroz_ · 2026-09-03
- Dev Ports True SPH Fluid Simulation to VRChat Udon, May Release on Booth — Michael_Moroz_ · 2026-09-03
- PRO-Step: Step-Level Process Reward Optimization Boosts Multi-Hop RAG (EMNLP 2026) — _reachsumit · 2026-09-03
- Meta's CORAL: An LLM-Native Harness That Continuously Optimizes Production Recommenders — _reachsumit · 2026-09-03