MMMMM: a taxonomy of multilingual multimodal misinformation built from seven languages of real data

copenlu · hf · 2026-09-01

Researchers released MMMMM, a unified taxonomy and dataset for studying multilingual multimodal misinformation. Image-text misinformation is more prevalent and harmful than its text-only counterpart, yet understudied due to lacking real-world-grounded taxonomies and scalable annotation tools.

The work proceeds in three steps: collecting a large-scale dataset of real misinformation from Twitter/X in seven languages; developing a comprehensive taxonomy via in-depth qualitative analysis; and operationalizing it through an automated VLM-based multi-step annotation pipeline with human validation.

Notable findings: AI-generated content is disproportionately prevalent in technology and science misinformation, while vaccination misinformation heavily reuses images from news outlets to assert credibility. The authors suggest mitigation should be applied strategically rather than uniformly.

Original post →

More from Safety

Safety channel →