Xiaohongshu’s UniNote unifies multimodal embedding and ranking in one model
小红书技术REDtech · wechat · 2026-07-24
Xiaohongshu’s tech team explains UniNote, a unified multimodal embedding model for retrieval and ranking.
- The system targets complex UGC notes that combine text, images, and OCR signals, aiming to solve industrial item-to-item retrieval more efficiently.
- Instead of a separate embedder plus reranker pipeline, UniNote uses a single multimodal model to do representation and ranking together.
- Training is split into two stages: first, SFT builds strong embeddings; then RL optimizes ranking quality with a structured reward design.
- The model also adopts Matryoshka Representation Learning so the same embedding can be truncated into multiple dimensions, from 64 to 4096, for different latency/quality trade-offs.
- In experiments, UniNote outperforms strong baselines across most retrieval tasks, though OCR-heavy tasks still leave room for improvement.
- The article ends with a hiring call for multimodal and content-understanding engineers.
More from Companies & People
- Enterprise AI agents still look like loops, memory, tools, and context stitched together — Johannascot · 2026-07-24
- Anthropic Says Claude Now Writes Over 80% of Its Merged Code — nathanbenaich · 2026-07-24
- StarRocks says the agent era needs a GPU-native database, not more CPU clusters — 智东西 · 2026-07-24
- AI is collapsing coordination costs, but firms may still outlast the market — prasanna_says · 2026-07-24
- Cohere moves ML Summer School to Twitch after 7,000+ signups — Cohere_Labs · 2026-07-24
- AI meeting notetakers hit a privacy wall as Teams adds an “Unverified” lobby — shashib · 2026-07-24