RecreationWorld: five-platform CUA benchmark — GPT-6 Astra hits 58.1% but passes all tests on just 2.8%
Shuai Bai · hf · 2026-09-21
A new Hugging Face release, RecreationWorld, provides scalable and verifiable environments for hybrid computer-use agents (CUAs) across Ubuntu, macOS, Windows, Android, and Web.
Design: Tasks are built around "recreation" — given a running reference app, the agent must discover its behavior and faithfully reimplement it with no prescribed workflow, mixing GUI exploration, coding, and visual verification. The running reference acts as an oracle for hidden behavioral tests, yielding execution-grounded rewards.
Training & transfer: Trajectories scaled from quality open-source apps; models trained on them improve across five out-of-distribution coding and hybrid CUA benchmarks and verify rendered outputs more often.
RecreationBench: 250 held-out tasks with programmatic and visual assertions, human-reviewed before freezing.
Results: GPT-6 Astra leads at 58.1% overall but passes all programmatic tests on only 2.8% of tasks. Agents reproduce static UI structure more reliably than interactions and computed outputs, and generated apps are smaller and more monolithic than references. Benchmark, environments, and test suites are open-sourced.
More from coding & agent
- KDE Drafts AI Policy: Use LLMs, But Don't Tell Anyone — carsonfarmer · 2026-09-21
- 9 practical Jev API recipes: browser agent flies through Google Flights in 7 seconds — alexcovo_eth · 2026-09-21
- GitHub's 17.4k-Star stop-slop Teaches LLMs to Strip AI Writing Tells — tom_doerr · 2026-09-21
- Solo attorney running agent fleet: evidence gates to stop coding-agent regression loops — Specialist_Call_1257 · 2026-09-21
- Code review classification saves 122k tokens, author says even 10% savings is a free win — draginol · 2026-09-21
- Microsoft open-sources IQ Solution Accelerator unifying enterprise data, knowledge and decisions — adnan_hashmi · 2026-09-21