NormViz benchmark: best model Gemini 3 Flash scores just 25.3% on visual cultural norms across 16 countries
StellaLisy · x · 2026-10-01
- A COLM 2026 paper introduces NormViz, a benchmark and framework measuring how well multimodal models understand visual cultural norms across 16 countries.
- Examples include gifting a plate with four koi in Japan (4 sounds like "death") and sticking chopsticks upright in rice in China (resembles funeral incense).
- The best model, Gemini 3 Flash, scores only 25.3%, showing top models largely fail at visual cultural taboos.
More from Models
- Cloudflare open-sources its first homegrown decision models clef and clef-flash — ritakozlov · 2026-10-01
- First time seeing a plan with unlimited human usage but capped agent usage — _philschmid · 2026-10-01
- NPR Deep Dive on the Murky State of Third-Party AI Model Evaluations — Miles_Brundage · 2026-10-01
- Simple interfaces for WIP models: a high-leverage habit for faster iteration — gowthami_s · 2026-10-01
- Dwarfstar's Bespoke Quants Run Qwen Fast on a 96GB M3 Ultra — TheRealJesus2 · 2026-10-01
- Codex lead says usage on the primary dot is virtually unlimited, new limits needed — taherdhanera · 2026-10-01