Dynamic multimodal fact-checking benchmarks still hide 17%–29% contamination risk

Haorui He · hf · 2026-07-29

What the paper shows

This paper revisits the common assumption that dynamic multimodal fact-checking benchmarks are contamination-free simply because their claims were published after an LLM’s knowledge cutoff.

Using both the state-of-the-art static benchmark AVeriTeC and a newly built dynamic benchmark, ClaimReview2025Q4, the authors find that:

They then re-evaluate SOTA LLMs under a stricter contamination-controlled protocol and provide practical guidelines for more trustworthy multimodal automated fact-checking evaluation.

Original post →

More from Multimodal

Multimodal channel →