ICAE-Bench tests coding agents as interactive project builders from fuzzy briefs
Zhongyuan Peng · hf · 2026-07-24
ICAE-Bench evaluates coding agents as interactive project builders
The paper argues that current coding-agent benchmarks lag behind the reality of vibe coding, where agents are expected to turn fuzzy product intent into working software rather than simply complete fully specified tasks.
It introduces ICAE-Bench, which starts from ambiguous requirements grounded in real open-source repositories with executable behavior. To make the setup realistic and reproducible, it adds:
- an automated User Agent that can reveal hidden constraints without inventing new requirements or leaking implementation details
- black-box tests plus multi-dimensional diagnostics
- metrics for functional correctness, semantic/API similarity, structural fidelity, design quality, and interaction quality
The goal is to evaluate agents in a dynamic, project-building setting instead of static task completion.
More from coding & agent
- Treating agents like 50 First Dates: a 3-layer context system so every conversation doesn't start from zero — evielync · 2026-09-11
- SmolVM open-sources persistent computer infrastructure for agents that outlive chat sessions — aniketmaurya · 2026-09-11
- ARRM targets silent economic regressions in AI agents that functional tests miss — Beautiful_Belt_601 · 2026-09-11
- Dev builds browser 3D pizza delivery game with Claude: physics, GPS pathfinding, traffic AI — vinishkapoor · 2026-09-11
- Build X Carousel Posts from One Wide Image: A Splitter Tool Plus YouMind Skill Workflow — sujingshen · 2026-09-11
- "Anyone still coding the old way?" The joke capturing post-AI programming culture — lxfater · 2026-09-11