TerraVis quantifies world-grounded visual consistency failures in text-to-image models
the-aiml · hf · 2026-10-09
TerraVis evaluates world-grounded visual consistency in text-to-image generation via an 18-type taxonomy of object-, interaction-, and scene-level violations, detected by multi-stage MLLM workflows. It achieves the strongest correlation with human judgments among existing metrics, and shows models strong on conventional metrics still exhibit substantial world-consistency failures. Code is public.
More from Multimodal
- Seedance 2.5 AI video demo stuns users with crazy detail level — SimplyAnnisa · 2026-10-09
- Claude Motion ships with day-one HyperFrames Studio integration for video editing — sean_t_strong · 2026-10-09
- Whistle: an open 16.9 MB speech-to-text model that runs on CPU with 11 ms first token in 7 languages — solyarisoftware · 2026-10-09
- Open-source huashu-art-motion turns coding agents into art-animation studios, 2.6k stars — AlchainHust · 2026-10-09
- VOCALOID7 AI Megpoid voicebank based on Megumi Nakajima opens preorders with 8 voice types — CurieuxExplorer · 2026-10-09
- Iris-3B open-sourced: a 3B pixel-space generation and general vision learner — Total-Resort-3120 · 2026-10-09