VISTA: First Benchmark for Figma-to-Web Coding Agents
机器之心 · wechat · 2026-07-06
A joint team from the University of Arizona, Zoom, and Stony Brook University introduced VISTA—the first end-to-end "visual spec to Web app" Coding Agent benchmark, accompanied by a continuously updated leaderboard. Unlike traditional SWE benchmarks that focus on fixing existing code, VISTA requires agents to build a fully functional, interactive Web application from scratch based on product requirements, web designs, and Figma files, evaluating across multiple dimensions including product quality, efficiency, and cost.
The leaderboard reveals that the Coding Agent competition has evolved from a "battle of models" to a "model + harness" systems competition. Leading models like fable-5, Claude Opus 4.8, GPT-5.5, and GLM-5.2 can now build complete Web apps, but the highest overall score remains below 0.3. "Best" doesn't mean "fastest and cheapest": the top-ranked fable-5 consumes an average of 750,000 tokens per task, whereas GLM-5.2 uses about 300,000 and GPT-5.5 about 280,000, highlighting distinctly different engineering styles among the models.
More from coding & agent
- Tweaked orchestration skill turns agents into self-policing workflow — pvncher · 2026-07-27
- A practical map of 11 protocols in the modern AI agent stack — TheTuringPost · 2026-07-27
- Qwen Code nightly adds Goal v3 orchestration and workspace channel controls — qwen-code-ci-bot · 2026-07-27
- NVIDIA says Nemotron 3 Ultra hit 97.1% on agentic RTL chip-design tasks — NVIDIAAI · 2026-07-27
- Tokyo Agent Forge hackathon shipped production-ready AI agents in one day — DavidBennett__ · 2026-07-27
- Long-running agents will need immutable event logs, this thread argues — sebpaquet · 2026-07-27