VIEScore with GPT-4v hits 0.3 Spearman vs humans — below the 0.45 human-human bar
le_james94 · x · 2026-09-22
- The author sets a credibility bar for image-generation judges: a judge model must agree with humans more often than humans agree with each other.
- Numbers: VIEScore with GPT-4v scores 0.3 Spearman against human ratings, while human-vs-human sits at 0.45 — the judge hasn't cleared it.
- Takeaway: until that bar is cleared and calibration is published, decomposed metrics remain preference scores, not objective evaluation.
Related event: Image Eval Reliability and the Fragility of C2PA Metadata(2 posts)→
More from Multimodal
- Higgsfield launches Genjutsu: video motion transfer and object swap tool — tisch_eins · 2026-09-22
- Luma's hybrid production makes underwater shoots far cheaper and easier — mrjonfinger · 2026-09-22
- Training Krea2 character LoRAs on a 4060 8GB in ~40 minutes — BigBullshitta · 2026-09-22
- VidRush launches Persona I, pricing long-form AI avatar video as low as $0.6/min — GCWebDesigner · 2026-09-22
- Tiny 100M text-to-image model runs on an Android phone via Termux — Tight_Commercial7 · 2026-09-22
- krea2-bbox-turbo Full-Rank Finetune Hits Epoch 14, Release Planned at Epoch 20 — Amazing_Painter_7692 · 2026-09-22