Tsinghua's KBMR embeds images by semantic entity, boosting knowledge-based visual QA
Tsinghua · hf · 2026-09-03
Tsinghua researchers introduced KBMR, a retrieval method for knowledge-based visual question answering.
- Instead of surface visual similarity, KBMR uses a multimodal language model to embed images by semantic identity, aligning embeddings with the underlying entity rather than appearance.
- Training relies on continuous distillation and hard negative sampling to sharpen retrieval.
- The authors report improvements in both retrieval quality and downstream VQA accuracy.
More from Multimodal
- MiniMax H3 video workflow: 8-step sampling, 2-step upscale, and a first-frame darkening fix — Major_Specific_23 · 2026-09-03
- Fable 5.1 generates cursive Chinese script — but at 0.25x speed the stroke order is all wrong — windx0303 · 2026-09-03
- Open-source "Yingzao" skill turns travel snapshots into magazine-grade architecture posters — op7418 · 2026-09-03
- 10-year brand designer shows how to build a full brand system from one reference image with AI — GCWebDesigner · 2026-09-03
- Minimax FL2VA H3 supports Reference Voice — demo with Japanese Goku voice — Willow-External · 2026-09-03
- Open-sourced an experimental standalone DLSS 5 video player for neural rendering — 2600th · 2026-09-03