Task prefixes for embedding: pair SearchQuery with Document for retrieval
tomaarsen · x · 2026-10-07
Text inputs to Google's new embedding model use task prefixes: SearchQuery, QuestionAnswering, CodeRetrieval, Classification, Clustering and more. For retrieval, pair the query with SearchQuery and corpus text with Document (which adds "title: none"—format real titles manually; no prefix for media). Dimension tradeoff: 128d gives 6x smaller vectors but MMEB v2 drops from 59.01 at 768d to 45.65; the card positions 128d for text-only, while 256d loses far less for mixed modalities.
More from Multimodal
- DJ app's MCP server lets Claude drive real synths and drums instead of generating audio — tech_sand · 2026-10-07
- Spira launches Maxima 1.0, a text-to-social-video model post-trained on trending data — JaynitMakwana · 2026-10-07
- Nano Banana 2.1 Finally Nails the Monstera Leaf Word-Gap Prompt — fofrAI · 2026-10-07
- ID-Forcing Keeps Long Video Generation In-Distribution, Enabling Minute-Scale Videos Without Fine-Tuning — _akhaliq · 2026-10-07
- Google's multimodal embedding: images cost 280 tokens, audio 25 tokens per second — tomaarsen · 2026-10-07
- Google's multimodal embedding: vision token budget tunable from 70 to 1,120 per frame — tomaarsen · 2026-10-07