VISTA: First Benchmark for Figma-to-Web Coding Agents

机器之心 · wechat · 2026-07-06

A joint team from the University of Arizona, Zoom, and Stony Brook University introduced VISTA—the first end-to-end "visual spec to Web app" Coding Agent benchmark, accompanied by a continuously updated leaderboard. Unlike traditional SWE benchmarks that focus on fixing existing code, VISTA requires agents to build a fully functional, interactive Web application from scratch based on product requirements, web designs, and Figma files, evaluating across multiple dimensions including product quality, efficiency, and cost.

The leaderboard reveals that the Coding Agent competition has evolved from a "battle of models" to a "model + harness" systems competition. Leading models like fable-5, Claude Opus 4.8, GPT-5.5, and GLM-5.2 can now build complete Web apps, but the highest overall score remains below 0.3. "Best" doesn't mean "fastest and cheapest": the top-ranked fable-5 consumes an average of 750,000 tokens per task, whereas GLM-5.2 uses about 300,000 and GPT-5.5 about 280,000, highlighting distinctly different engineering styles among the models.

Original post →

More from coding & agent

coding & agent channel →