Evaluating image editing models with a single score hides the nuance behind a lower score. EdiVal-Agent (ICLR 2026) splits evaluation across instruction following, content consistency, and visual quality. Its agentic judge reached 81.3% agreement with human judgments, compared with 75.2% for a VLM-only evaluator and 68.9% for CLIP dir. ↧ Closing the loop: agentic evaluation for image editing foundation m... EdiVal-Agent uses agentic AI to evaluate image editing models turn-by-turn, revealing instruction-following failures a single score would miss.