Autoregressive Retriever (ARR) Refines Queries with Retrieved Item Feedback via SFT and RL
_reachsumit · x · 2026-10-09
Paper [2610.11666] introduces ARR, a multimodal retriever that addresses the limitation of encoding a query once and ranking items independently.
- Core idea: alternate between retrieving an item and updating the query embedding with its content, then rank the collection with the final embedding
- Training: SFT with stepwise contrastive supervision teaches the encoder to use feedback; RL treats feedback items as actions, optimizing selection by the final reciprocal rank of relevant items
- A query-side adapter enables optimization against a fixed item index
- Results: outperforms baselines on in-domain and zero-shot benchmarks; feedback helps at inference, and training with feedback also improves the initial query embedding
More from Multimodal
- You can now spot Opus AI video slop by its sound: synced beats as a fingerprint — hudzah · 2026-10-09
- Monkey King riding a tiger: AI video nails a stunning Chinese-style action scene — lucky-plume · 2026-10-09
- Sony's Syn-Omni: Shared + Expert LoRA Paths Beat Omnimodal Embedding Baselines Across 81 Tasks — _reachsumit · 2026-10-09
- LEGO: lifting-free exocentric-to-egocentric video generation beats depth-lifting SOTA pipelines — 25frms · 2026-10-09
- RISEBench++: 65 reasoning-based visual editing tasks; best model GPT-Image-2.5 hits only 56.6% — VisionXLab · 2026-10-09
- VibeEdit Replaces Text Prompts with Canvas Marks, Scoring 79.9 on Edit Benchmark — Sydney-Uni · 2026-10-09