Tri-PvP benchmark exposes visual bias in omni-modal LLMs via evidence conflicts
Yen-Ting Piao · hf · 2026-09-23
Tri-PvP is a new benchmark that measures how omni-modal LLMs handle conflicts between perceptual and propositional evidence. It reveals pronounced visual bias and asymmetric preferences for evidence forms, and shows modality effects are decodable in early layers and resistant to surface-level mitigation, suggesting unreliable evidence integration under conflict.
More from Research
- Pokemon benchmark Paradigm 3: Astra generalizes to scrambled maps and fan-made games while rivals memorize — gleech · 2026-09-24
- AI Completes Fan-Made Pokemon Brown in 10K Steps: Real Generalization or Whack-a-Mole? — gleech · 2026-09-24
- Astra Beats Fan-Made Pokemon Brown in ~10K Steps Without Memorizing It — gleech · 2026-09-24
- Genome language model Omnii enters research preview, unveiled on Latent Space podcast — exnx · 2026-09-24
- Economist Novosad: AI detector debates trip over FPR vs FNR confusion — paulnovosad · 2026-09-24
- exNX unveils research preview of genome language model Omnii on Latent Space — exnx · 2026-09-24