Lambda's EdiVal-Agent uses an agentic judge for image editing evals, hitting 81.3% human agreement

TheZachMueller · x · 2026-09-18

Lambda, with UT Austin, UCLA, and Microsoft, released EdiVal-Agent (ICLR 2026), an open framework that turns evaluation of image editing foundation models into an agentic workflow.

Why it matters: A good edit must follow the instruction exactly, preserve everything else, and keep visual quality — multi-turn editing adds more constraints. Single scores hide this nuance, human review doesn't scale, and classic metrics like CLIP fall short.

Approach and results: EdiVal-Agent splits evaluation into instruction following, content consistency, and visual quality, judged by an agentic pipeline:

It closes the train→evaluate→improve loop for image editing models with scalable automated evaluation.

Original post →

More from Multimodal

Multimodal channel →