Qwen releases RecreationBench: 250 tasks testing agents that rebuild real apps
burny_tech · x · 2026-09-21
Qwen released RecreationBench on Hugging Face: a 250-task benchmark spanning Ubuntu, macOS, Windows, Android and Web, designed to test hybrid computer-use agents that recreate real applications from scratch.
Unlike screenshot-and-click computer-use evals, it measures an agent's full understanding and reimplementation of real apps, offering a new yardstick for end-to-end software-building agents.
More from coding & agent
- SpaceXAI new hire's week 1: intense pace, fixes Grok Bot browser bug, says agents have far to go — elonmusk · 2026-09-21
- Developer tells agent to 'figure out' the API key itself — and it does — shakoistsLog · 2026-09-21
- Pretty Mermaid: Open-Source Skill Renders Mermaid to SVG/PNG/ASCII Locally, No Browser — tom_doerr · 2026-09-21
- Handing my tax credentials to an AI agent: great UX, unsettling security — zeeg · 2026-09-21
- Coding agent AdaL opens free access to 10 users, inviting harshest criticism — Zachly · 2026-09-21
- Dev spins up a custom Grok Bot "anew" to serve free AI webpages on X — round · 2026-09-21