SetwiseEvalKit scores document sets, not just single search results
Kailin Jiang · hf · 2026-07-23
Why this matters
This paper argues that once LLMs and agents consume search results directly, the quality of the document set matters more than single-document relevance.
- The authors show that existing evaluation tools miss inter-document effects such as redundancy, conflict, and complementarity.
- They introduce SetwiseEvalKit, a three-level, nine-dimension benchmark with about 28K evaluation rubrics.
- Across 12 rerankers, even the best method covers no more than 45% of the rubric space, and cross-document coordination remains weak.
- They then propose Rubric4Setwise, a training-free method that turns rubric criteria into document-set selection signals.
- Rubric4Setwise achieves the best downstream generation performance with fewer documents and search rounds, and is the only method that stays state of the art across both short-form and long-form settings.
More from Research
- MICCAI FLARE 2026 asks if one AutoML system can handle segmentation and classification — yuyinzhou_cs · 2026-07-23
- Pan-cancer CT segmentation dataset adds 17,000 labeled cases and a new challenge — yuyinzhou_cs · 2026-07-23
- SIGIR 2026 paper asks whether QPP can pick the best query variant before RAG costs kick in — mrdrozdov · 2026-07-23
- Kepler v0.1 trains robot vision, touch, and pose into one shared world model — freelerobot · 2026-07-23
- GPT-5.5 helps generate five new Banach space results in a math discovery study — ChrSzegedy · 2026-07-23
- One in seven MCP registry source repos is no longer publicly reachable — mcpindex · 2026-07-23