Trident: Multi-Aspect Page Annotation for Long-Document VQA

_reachsumit · x · 2026-08-18

Addressing the bottleneck in long-document multimodal VQA where rerankers struggle with visual evidence, the paper proposes Trident. Trident-R converts candidates into LLM-readable semantic records with visual captions and structure tags, enabling text-only LLM rerankers to select relevant evidence. The method substantially improves retrieval F1 across heterogeneous pools.

Original post →

More from Research

Research channel →