Meta's Guided SID: Pinning Coarse Semantic ID Levels Makes Fine Levels Learnable
_reachsumit · x · 2026-09-22
Meta published Guided SID on arXiv (2609.22227) for generative retrieval, where items are represented by short Semantic IDs and recommendation is cast as autoregressive generation. Since tokenizers trained independently don't align with downstream LLMs or the end task, existing SID systems need extra bridging effort.
Method: force the coarse RQ-VAE levels to encode a predefined categorical attribute — text-grounded (legible to the LLM) and task-relevant — via deterministic supervised index assignment while keeping codebooks learnable; a trie-merge construction maps high-cardinality attributes onto the fixed code budget.
Results: guiding costs nothing intrinsically; in a matched end-to-end A/B, the guided retriever improves recall@k at every list length (up to 1.36 points).
More from Research
- Boltzbit previews paper claiming BAST lets LLMs learn up to 1,000x faster than SOTA training — jmhernandez233 · 2026-09-22
- MiMo's GRS and GAR: scoring what makes an RL answer genuinely good, not just passing — tokenbender · 2026-09-22
- MiMo v2.6 details: grounded environment synthesis and multi-harness training for open source users — tokenbender · 2026-09-22
- MiMo leans on CodeMidas-style source-driven synthesis and agent-driven long-horizon tasks — tokenbender · 2026-09-22
- Xiaomi MiMo v2.6 ships with RL as the hero, scaling batch size, env diversity and grader compute — tokenbender · 2026-09-22
- Glasshouse v0.1 launches: an open memory benchmark with 2,847 questions over a 1.97M-token conversation — True_Mongoose_7073 · 2026-09-22