Grok 4.6 tops VISTA benchmark, turning Figma designs into web apps at $2.38 per task
XFreeze · x · 2026-08-17
Grok 4.6 has topped the VISTA benchmark, which tests AI coding agents' ability to understand Figma designs and build functional web apps. It outperformed Claude Fable 5, Opus 5, and GPT-5.6 Sol, with a cost of about $2.38 per task. This demonstrates the full design-to-code capability, which is practically relevant for developers.
More from coding & agent
- LLM Agents Prone to Reward Hacking in Code Verification — TuhinChakr · 2026-08-17
- Using Claude to Analyze 10k Support Emails for Feature Requests — mhmazur · 2026-08-17
- "Assume Author is an Idiot": Prompt Engineering Stricter Code Agents — aronchick · 2026-08-17
- audio.cpp 0.6 Released: Adds MiniMax-H3 Text-to-Audio, 5 New Model Families — Acceptable-Cycle4645 · 2026-08-17
- Multi-Agent Workflow: Grok Bot Playtests and Fixes Cursor-Built Game — DeryaTR_ · 2026-08-17
- AI agents make coding non-linear: idle waits vs overload bursts — StewartalsopIII · 2026-08-17