RecreationWorld: five-platform CUA benchmark — GPT-6 Astra hits 58.1% but passes all tests on just 2.8%

Shuai Bai · hf · 2026-09-21

A new Hugging Face release, RecreationWorld, provides scalable and verifiable environments for hybrid computer-use agents (CUAs) across Ubuntu, macOS, Windows, Android, and Web.

Design: Tasks are built around "recreation" — given a running reference app, the agent must discover its behavior and faithfully reimplement it with no prescribed workflow, mixing GUI exploration, coding, and visual verification. The running reference acts as an oracle for hidden behavioral tests, yielding execution-grounded rewards.

Training & transfer: Trajectories scaled from quality open-source apps; models trained on them improve across five out-of-distribution coding and hybrid CUA benchmarks and verify rendered outputs more often.

RecreationBench: 250 held-out tasks with programmatic and visual assertions, human-reviewed before freezing.

Results: GPT-6 Astra leads at 58.1% overall but passes all programmatic tests on only 2.8% of tasks. Agents reproduce static UI structure more reliably than interactions and computed outputs, and generated apps are smaller and more monolithic than references. Benchmark, environments, and test suites are open-sourced.

Original post →

More from coding & agent

coding & agent channel →