EVIE: multimodal-judge-trained visual document retrievers with 6 nested embedding sizes per checkpoint
_reachsumit · x · 2026-10-09
Tencent researchers introduce EVIE, a family of native visual document retrievers with three key innovations:
- Evidence-judged data governance: a multimodal judge relabels answer-bearing positives and filters unreliable negatives.
- Bidirectional teacher-student learning: symmetric listwise distillation plus prefix-based Matryoshka representation learning (Prefix-MRL), so one checkpoint serves six nested embedding dimensions without re-encoding.
- Hierarchical agglomerative index compression (HAC): clusters page tokens with spatial regularization and stores semantic centroids.
EVIE targets VDR pain points: OCR latency and lost visual cues, single-vector granularity limits, and oversized multi-vector indexes.
More from Research
- RSI-Exam lands in State of AI report: top model Opus 5.5 scores just ~0.53 — cihangxie · 2026-10-09
- KAIST's ME-World tackles multi-agent egocentric world modeling with joint denoising — kaist-ai · 2026-10-09
- Controlled study: LLMs struggle to recover latent sequential structure despite long context — illinois · 2026-10-09
- TerraVis quantifies world-grounded visual consistency failures in text-to-image models — the-aiml · 2026-10-09
- OneSearch-VL: unified multimodal deep research agent beats Qwen3-VL by 20 points — Hongyu Li · 2026-10-09
- SpaceCast-Bench: best VLM hits 58.0% on predictive spatial reasoning vs 87.2% human — zju · 2026-10-09