Tencent Proposes RubricRanker: Training Rerankers for Deep Research Agents
_reachsumit · x · 2026-08-05
Existing retrieval systems generally select documents based on relevance matching. However, for deep research agents, individually well-matched documents may not form a high-quality set that satisfies complex query requirements (e.g., diversity, conciseness, and authority).
To address this, Tencent introduced search-oriented rubrics that explicitly define the criteria for high-quality document sets. Based on these hierarchically structured rubrics synthesized by an LLM, the researchers trained a document reranker model named RubricRanker.
The training framework consists of two stages: rubrics-guided supervised fine-tuning (SFT) and rubric-based reinforcement learning (RL). Experiments demonstrate that RubricRanker outperforms the strongest baseline by 2.6 points on four deep research benchmarks and generalizes well to five RAG benchmarks.
More from coding & agent
- Jev-style models called a major unlock for fast, cheap browser agents — multiply_matrix · 2026-09-22
- Generating responsive UIs in under 2 seconds with shadcn plus Mobbin MCP — msharmas · 2026-09-22
- Nat Friedman admits Meta's Muse was inspired by OpenClaw, bought hundreds of Mac minis for the team — EdwardSun0909 · 2026-09-22
- SkillLift cuts agent skill-evolution token cost 40-70% by ranking, not rollouts — dair_ai · 2026-09-22
- Deel launches Akai: 8,000 agents doing work of ~600 employees, added $140M ARR in 90 days — AIwithGhotai · 2026-09-22
- exe.dev wins over developers: SSH into root VMs, plus the underrated Shelley coding agent — davidcrawshaw · 2026-09-22