Engineer's eval: cross-encoder rerankers lift 30% of cases but wreck the retriever in another 30%
bo_wangbo · x · 2026-10-08
Engineer bowangbo shares first-hand eval results on cross-encoder rerankers: on his eval sets, they improved 30% of cases, did nothing in 40%, and ruined the first-stage retriever in another 30%, leaving him little confidence in cross-encoders for OOD tasks. The thread also asks whether late-interaction models count as cross-encoders.
Related event: Debate: Are Cross-Encoder Rerankers Dead?(4 posts)→
More from Models
- Anthropic red team: GLM-5.3 safeguards bypassed 64%-100% in simulated cyber tests — dl_weekly · 2026-10-09
- d1-omni-600M sorts your voice notes in ~160ms, fully in-browser on WebGPU — iamrobotbear · 2026-10-09
- Arena Launches Alignment Index: 90K Real Agent Sessions Rank GPT-6.1-Sol Safest at 87.9 — arena · 2026-10-09
- Dan Shipper warns of the "agent death spiral" that burns millions of tokens — every · 2026-10-09
- Gemini 4 Argon Ties GPT-6 Astra at 53 on AA Index at ~60% of the Cost — DeepLearningAI · 2026-10-09
- MTP in llama.cpp now rivals ds4: GLM 5.3 Flash sets new Apple Silicon decode record — challis88ocarina · 2026-10-08