VIEScore2 unifies image eval with spatially grounded defect localization, beating Gemini-3-Flash
TIGER-Lab · hf · 2026-10-08
TIGER-Lab introduces VIEScore2, a unified evaluator for image generation and editing that predicts quality scores together with defect locations rather than just a scalar score.
Details:
- Represents images as an N×N grid and jointly predicts scores and defect locations in one pass; the text-native grid representation unifies heterogeneous spatial supervision and enables verifiable post-training objectives;
- Trained on 38K examples spanning score-only, localization-only, and joint supervision across generation and editing;
- Applies GRPO after SFT with rewards combining cell-level Dice overlap, score accuracy, and format validity; a parameter-free parser turns predictions into readable explanations.
Results: overall-score SRCC of 0.601 vs 0.491 for Gemini-3-Flash, the strongest zero-shot general-purpose VLM baseline; outperforms general VLMs and specialized spatial evaluators on three of six localization benchmarks, top-three on five including out-of-training datasets.
Related event: TIGER-Lab Releases VIEScore2, a Unified Evaluator for Image Generation(2 posts)→
More from Multimodal
- KAIST's Tetris3D Reconstructs 3D Scenes With Physically Coherent Objects, Plus 1.2M-Scene ComOb Dataset — kaist-ai · 2026-10-08
- Suno Praised as Good Enough to Eventually Overtake Spotify — petergyang · 2026-10-08
- 3.5M-param anime upscaler applies Looped-DiT recurrence, gains +0.6dB in 4 loops — NobodySnJake · 2026-10-08
- HKU and VAST's Mira-Scene fixes object placement in single-image 3D scene reconstruction — jiqizhixin · 2026-10-08
- Musk Hypes Grok Bot After User Generates Full Explainer Video From a 15-Second Prompt — elonmusk · 2026-10-08
- AI pop singer Claudia open-sourced: downloadable soul via MCP server, skills and character sheets — Promptmethus · 2026-10-08