Dynamic multimodal fact-checking benchmarks still hide 17%–29% contamination risk
Haorui He · hf · 2026-07-29
What the paper shows
This paper revisits the common assumption that dynamic multimodal fact-checking benchmarks are contamination-free simply because their claims were published after an LLM’s knowledge cutoff.
Using both the state-of-the-art static benchmark AVeriTeC and a newly built dynamic benchmark, ClaimReview2025Q4, the authors find that:
- 17.09%–29.30% of post-cutoff claims may still be contaminated.
- Many supposedly novel claims can still be verified from public knowledge that existed before the cutoff, either directly or by combining multiple sources.
- Contamination can meaningfully distort evaluation, inflating Macro-F1 by up to 11.34 points and even changing system rankings.
They then re-evaluate SOTA LLMs under a stricter contamination-controlled protocol and provide practical guidelines for more trustworthy multimodal automated fact-checking evaluation.
More from Multimodal
- MODUS turns a decoder-only model into a single any-to-any multimodal system — EPFL-VILAB · 2026-07-29
- GPT Image 2 storyboard packs a pilot’s tea break and a mecha battle into one 3x3 grid — Practical_Low29 · 2026-07-29
- Prompt tweak turns a generic scene into a Saudi heritage image with traditional details — aziz4ai · 2026-07-29
- Twelve Labs unveils a video intelligence stack built around search, memory and agentic workflows — qdrant_engine · 2026-07-29
- A colleague built a silly AI skill that turns any song into a parody music video — oran_ge · 2026-07-29
- Magnific brings its AI suite natively into Photoshop, Premiere and Figma — aziz4ai · 2026-07-29