Tsinghua ACMMM paper traces short-answer MLLM hallucinations to visual features, not language priors

新智元 · wechat · 2026-09-20

A Tsinghua team (Zhu Jun, Hu Xiaolin et al.) at ACMMM 2026 identifies "visual-origin hallucination": when MLLMs answer Yes/No, object hallucinations stem from visual feature extraction, not language priors.

Findings

Method

Results

Original post →

More from Research

Research channel →