Image Eval Reliability and the Fragility of C2PA Metadata
VIEScore with GPT-4v correlates only 0.3 with human ratings—below human-human agreement of 0.45—while lectures reveal FLUX.2 and Qwen-Image share a recipe and C2PA metadata can be wiped with a screenshot.
2026-09-22 ~ 2026-09-22 · 2 related posts
- Lecture: FLUX.2 and Qwen-Image share one recipe; C2PA metadata dies to a screenshot — le_james94 · 2026-09-22
- VIEScore with GPT-4v hits 0.3 Spearman vs humans — below the 0.45 human-human bar — le_james94 · 2026-09-22