ACuRL: zero-human-data continual learning for computer-use agents lands at NeurIPS

ysu_nlp · x · 2026-09-26

Author Tianci Xue announced that ACuRL (Autonomous Curriculum Reinforcement Learning) has been accepted to NeurIPS. The framework lets computer-use agents continually adapt to dynamic software environments with zero human annotation data.

Key components:

An intriguing finding: replacing the task generator or evaluator with the policy model itself still yields improvements, hinting at potential recursive self-improvement—unverified at scale due to compute limits. Full infrastructure (orchestrating hundreds of Linux environments) is open-sourced.

Original post →

More from coding & agent

coding & agent channel →