ICAE-Bench tests coding agents as interactive project builders from fuzzy briefs

Zhongyuan Peng · hf · 2026-07-24

ICAE-Bench evaluates coding agents as interactive project builders

The paper argues that current coding-agent benchmarks lag behind the reality of vibe coding, where agents are expected to turn fuzzy product intent into working software rather than simply complete fully specified tasks.

It introduces ICAE-Bench, which starts from ambiguous requirements grounded in real open-source repositories with executable behavior. To make the setup realistic and reproducible, it adds:

The goal is to evaluate agents in a dynamic, project-building setting instead of static task completion.

Original post →

More from coding & agent

coding & agent channel →