VISTA: A Benchmark for Web App Generation Agents
jiqizhixin · x · 2026-07-14
Researchers introduce **VISTA**, a benchmark evaluating whether LLM/agent can generate **functional and visually consistent** web applications based on rough specifications. ### Key Highlights - Data and input conditions cover 5 scenarios: text, screenshots, Figma snippets, and other varied spec inputs. - Evaluation goes beyond final code by integrating **DOM matching, browser testing, and CLIP** to measure structure, behavior, and visual alignment respectively. - Results show only a partial correlation between **visual fidelity** and **functional correctness**; agents' editing styles vary widely but don't significantly impact task quality. The project also provides a paper, GitHub, and project page.
More from coding & agent
- Autoresearch proposes packaging ML runs as studies with questions, analysis, and code diffs — morgymcg · 2026-07-21
- CHAP defines approvals, handoffs, and audit logs for human-agent workflows — DeliveryTechnical199 · 2026-07-21
- The author says Codex reached 20x and is now debugging spec decoding on a hybrid parallel setup — TheZachMueller · 2026-07-21
- Axcess adds an MCP connector for WCAG accessibility checks that scanners miss — modelcontextprotocol · 2026-07-21
- X post asks whether Cursor Composer, built on Kimi models, would also be banned — max_paperclips · 2026-07-21
- A developer’s Codex usage is draining pooled enterprise credits at a small company — Distinct_Relation_62 · 2026-07-21