VISTA: First Benchmark for Figma-to-Web Coding Agents
机器之心 · wechat · 2026-07-06
A joint team from the University of Arizona, Zoom, and Stony Brook University introduced VISTA—the first end-to-end "visual spec to Web app" Coding Agent benchmark, accompanied by a continuously updated leaderboard. Unlike traditional SWE benchmarks that focus on fixing existing code, VISTA requires agents to build a fully functional, interactive Web application from scratch based on product requirements, web designs, and Figma files, evaluating across multiple dimensions including product quality, efficiency, and cost.
The leaderboard reveals that the Coding Agent competition has evolved from a "battle of models" to a "model + harness" systems competition. Leading models like fable-5, Claude Opus 4.8, GPT-5.5, and GLM-5.2 can now build complete Web apps, but the highest overall score remains below 0.3. "Best" doesn't mean "fastest and cheapest": the top-ranked fable-5 consumes an average of 750,000 tokens per task, whereas GLM-5.2 uses about 300,000 and GPT-5.5 about 280,000, highlighting distinctly different engineering styles among the models.
More from coding & agent
- Anthropic researcher: 99% of engineers now run swarms of 300+ self-improving agents — AlishaOutridge · 2026-09-11
- Gergely Orosz: Shipping 10x PRs With AI Agents, Sites Fill With Small Regressions — ducha_aiki · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Astra storyboards plus Minimax H3 per-shot generation boost video success rates — Hailuo_AI · 2026-09-11
- Codex tip: use Sol with Astra and Luna sub-agents to save usage — pvncher · 2026-09-11
- agents-best-practices: a provider-neutral Agent Skill for designing and auditing agentic harnesses — tom_doerr · 2026-09-11