Alibaba's Qwen team launches RecreationBench to test hybrid computer-use agents by app recreation
TianbaoX · x · 2026-09-22
Alibaba's Qwen team introduced RecreationWorld / RecreationBench, a scalable, verifiable environment for evaluating hybrid computer-use agents that dynamically interleave CLI and GUI actions.
- Core idea: an agent must recreate an application purely by using it—testing exploration and execution—with programmatically verifiable scoring.
- Scale: 250 held-out applications across five native platforms (Ubuntu, macOS, Windows, Android, Web), scored on Prog (functional) and Visual channels.
- Leaderboard: the site compares models on Prog/Visual scores, threshold coverage (Prog ≥90% and =100%), and estimated per-task API cost, including models like GPT-6 Astra across platforms.
- The paper adds analysis of cost-vs-score tradeoffs, exploration behavior, and failure boundaries; code and website are public.
Related event: Alibaba's Qwen Team Releases RecreationBench for Hybrid Computer-Use Agents(2 posts)→
More from coding & agent
- Jev-style models called a major unlock for fast, cheap browser agents — multiply_matrix · 2026-09-22
- Generating responsive UIs in under 2 seconds with shadcn plus Mobbin MCP — msharmas · 2026-09-22
- Nat Friedman admits Meta's Muse was inspired by OpenClaw, bought hundreds of Mac minis for the team — EdwardSun0909 · 2026-09-22
- SkillLift cuts agent skill-evolution token cost 40-70% by ranking, not rollouts — dair_ai · 2026-09-22
- Deel launches Akai: 8,000 agents doing work of ~600 employees, added $140M ARR in 90 days — AIwithGhotai · 2026-09-22
- exe.dev wins over developers: SSH into root VMs, plus the underrated Shelley coding agent — davidcrawshaw · 2026-09-22