TerraVis quantifies world-grounded visual consistency failures in text-to-image models

the-aiml · hf · 2026-10-09

TerraVis evaluates world-grounded visual consistency in text-to-image generation via an 18-type taxonomy of object-, interaction-, and scene-level violations, detected by multi-stage MLLM workflows. It achieves the strongest correlation with human judgments among existing metrics, and shows models strong on conventional metrics still exhibit substantial world-consistency failures. Code is public.

Original post →

More from Multimodal

Multimodal channel →