UI2App shows screenshot fidelity still lags real interaction recovery
Grace Man Chen · hf · 2026-07-22
UI2App benchmark shows screenshot fidelity does not imply interaction recovery
UI2App is a new benchmark for visual interaction inference in executable web app generation: given screenshots alone, can a model recover the app’s behavior without any textual or behavioral hints?
Benchmark design
- 327 screenshots grouped into 45 state-coherent screenshot sets.
- Runnable multi-route web applications as targets.
- Evaluation across four dimensions: executability, navigation reachability, visual fidelity, and interaction inference.
- A new metric, IIS, scores whether the model correctly infers functional interactions and state management, while crediting any valid implementation rather than forcing a single reference solution.
Key result
- Across six frontier vision-language models, the best visual-fidelity model scores only 7.5 on IIS and ranks fourth.
- The IIS leader beats it by 5.2×.
- High-complexity interactions, especially cross-page state, remain a major bottleneck, with half of the models scoring zero on that dimension.
The benchmark suggests that static screenshot reconstruction is still far from capturing full application behavior.
More from Research
- The Thimble and the Waterfall: AI's Data Bottleneck and Feedback Loops — dyamins · 2026-07-22
- NVIDIA shows 22 SIGGRAPH papers and Omniverse tools for robot simulation — facontidavide · 2026-07-22
- Building a Knowledge Graph Without a Graph DB: 1000x Cheaper Than GraphRAG — TheRedfather · 2026-07-22
- Agentic RAG survey maps planner, retriever and refinement agents for complex retrieval — blaizedsouza · 2026-07-22
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22