Lambda's EdiVal-Agent uses an agentic judge for image editing evals, hitting 81.3% human agreement
TheZachMueller · x · 2026-09-18
Lambda, with UT Austin, UCLA, and Microsoft, released EdiVal-Agent (ICLR 2026), an open framework that turns evaluation of image editing foundation models into an agentic workflow.
Why it matters: A good edit must follow the instruction exactly, preserve everything else, and keep visual quality — multi-turn editing adds more constraints. Single scores hide this nuance, human review doesn't scale, and classic metrics like CLIP fall short.
Approach and results: EdiVal-Agent splits evaluation into instruction following, content consistency, and visual quality, judged by an agentic pipeline:
- Agentic judge: 81.3% agreement with human judgments
- VLM-only evaluator: 75.2%
- CLIPdir: 68.9%
It closes the train→evaluate→improve loop for image editing models with scalable automated evaluation.
More from Multimodal
- Arabic AI creator hails new Pika platform's model lineup and API access — Kyrannio · 2026-09-18
- Seedance 2.5 + GPT Image 2.5 combined in new Pika, creator shares prompt — Kyrannio · 2026-09-18
- FAST H3 V2 vs 3-Step LoRA: community tests weigh speed and quality in H3 video workflows — Strange_Limit_9595 · 2026-09-18
- EasyAI: open-source desktop GUI wraps ComfyUI into a one-click beginner workflow — garionhk · 2026-09-18
- Pika's New Platform Hands Creators an All-in-One AI Anime Studio — Kyrannio · 2026-09-18
- Game reportedly making $30M/month pitched as replicable with GPT agents plus Higgsfield API — xiaohu · 2026-09-18