Study Finds Embedding Models Often Ignore Retrieval Instructions; Distractor Fine-tuning Helps
_reachsumit · x · 2026-10-08
Researchers from the University of Turku released an arXiv paper, Your Prompt Should Do More, examining how prompted embedding models process retrieval instructions in asymmetric retrieval.
- Key finding: when the candidate pool includes distractor texts similar to the query, embedding models often fail to follow even simple task instructions — the detailed retrieval prompt is effectively ignored.
- Mechanism: the authors attribute this to current training setups and evaluation protocols, which lack query-side distractors and thus never force models to actually use the instructions.
- Fix: fine-tuning with added query-side distractors yields substantial improvements while minimally affecting other tasks.
Authors: Amanda Myntti, Jenna Kanerva, Veronika Laippala, Filip Ginter.
More from Models
- User claims Qwen3.8-Flash-Next-NVFP4 sounds uncannily like Opus 4.5 — natesiggard · 2026-10-08
- Dev endorsement: DeepSeek is the best bang-for-buck model, DSH an excellent harness — sull · 2026-10-08
- Creator's personal-assistant agent comparison: grok bot clearly best, Muse decent, Dots needs work — elonmusk · 2026-10-08
- Claude Haiku 5.5 tops highlighted benchmark scores, likely far smaller than GLM 5.3 Flash — rickasaurus · 2026-10-08
- User loses 3 years of Gemini chats; Google support claims they were 'never saved' — gtboy1994 · 2026-10-08
- Google's new embedding model is now live on mjpro — ciguleva · 2026-10-08