ACuRL: zero-human-data continual learning for computer-use agents lands at NeurIPS
ysu_nlp · x · 2026-09-26
Author Tianci Xue announced that ACuRL (Autonomous Curriculum Reinforcement Learning) has been accepted to NeurIPS. The framework lets computer-use agents continually adapt to dynamic software environments with zero human annotation data.
Key components:
- Autonomous exploration: the agent gathers initial experience on its own
- Curriculum-driven task generation from that experience
- CUAJudge, an automatic evaluator with 93% human agreement
An intriguing finding: replacing the task generator or evaluator with the policy model itself still yields improvements, hinting at potential recursive self-improvement—unverified at scale due to compute limits. Full infrastructure (orchestrating hundreds of Linux environments) is open-sourced.
More from coding & agent
- ambion 0.3.0 ships git over SSH on remote workstations and a two-model grading simulator — andreisavu · 2026-09-26
- Ambion 0.3.0 ships a collaboration kernel where agent work outlives its activation — andreisavu · 2026-09-26
- Memory backups may resurrect revoked agent permissions across AIs — tallmetommy · 2026-09-26
- Dev builds Spatial Composition on Krea Agent for deliberate image composition control — angrypenguinPNG · 2026-09-26
- Jev-Mem: System-One/Two-Inspired Memory Architecture Builds Agent Memory 6.6x Faster — omarsar0 · 2026-09-26
- Managing Claude Agents From Cursor via Herdr: A Developer's Workflow Worth a Look — letandrewcook · 2026-09-26