Attention Relay makes embedding models instruction-aware without training via LLM attention weights
_reachsumit · x · 2026-10-06
An arXiv paper proposes Attention Relay, a training-free method that passes an instruction-tuned LLM's attention weights into an embedder's attention pooling, transferring instruction-following ability to text embeddings.
- Works across 6 instruction-tuned LLMs (Qwen3, Llama 3.1, OLMo 3) and 10 embedding models with different tokenizers, sizes, and pooling types
- Analysis shows the LLM's later-layer attention tracks the instruction, largely due to instruction tuning, and relaying it makes the instructed aspect dominant in the embedding
More from Models
- Reflection's Billion-Dollar US Open-Weight Model Underperforms Every Major Chinese Model, Mocked Online — npinto · 2026-10-06
- Opus 5.5 uses fewer tokens than GPT-6.1 Sol while scoring better, fan argues — Angaisb_ · 2026-10-06
- OpenWork, an open-source Claude Cowork alternative, hits 45K downloads in two days — alex_verem · 2026-10-06
- Abacus.AI CEO: Chinese open-source models beat US labs on easy tasks, DeepSeek cheapest for agents — bindureddy · 2026-10-06
- HF researcher disputes GLiDE's Decision Index win over Jev as reasoning-boosted — antoine_chaffin · 2026-10-06
- ChatGPT quietly adds lifetime usage tracking to profile settings — Bpelks · 2026-10-06