VISTA: A Benchmark for Web App Generation Agents

jiqizhixin · x · 2026-07-14

Researchers introduce **VISTA**, a benchmark evaluating whether LLM/agent can generate **functional and visually consistent** web applications based on rough specifications. ### Key Highlights - Data and input conditions cover 5 scenarios: text, screenshots, Figma snippets, and other varied spec inputs. - Evaluation goes beyond final code by integrating **DOM matching, browser testing, and CLIP** to measure structure, behavior, and visual alignment respectively. - Results show only a partial correlation between **visual fidelity** and **functional correctness**; agents' editing styles vary widely but don't significantly impact task quality. The project also provides a paper, GitHub, and project page.

Original post →

More from coding & agent

coding & agent channel →