AutoWorldModel-Bench: A New Benchmark for Autonomous Coding Agents

Marjan Moodi · hf · 2026-08-13

AutoWorldModel-Bench is a state-centric benchmark. It evaluates autonomous coding agents on open-ended world-model research by having them iteratively improve a starter model across game environments using a shared structured-state format.

Original post →

More from coding & agent

coding & agent channel →