CompoWorld scales general agent training by composing environments from reusable services

AllSpark-Research · hf · 2026-09-29

CompoWorld introduces Compositional Environment Scaling, expanding the agent task space by composing a finite library of reusable services instead of generating tasks within a single environment.

Coding agents turn tool specs into verified services with typed states and shared interfaces; a world model stands in for tools that can't be reliably implemented. A random-walk procedure connects services via dependency graphs, generating and verifying tasks that require cross-service information flow. Verified trajectories feed SFT, while a Completion-Focused Rubric Reward steers RL toward full task completion by emphasizing lower-pass-rate criteria within each rollout group.

The team built 448 services exposing 10,130 tools and trained Qwen3.6-35B-A3B with 3K SFT trajectories and 1K RL tasks. The model improves its backbone by an average of 9.17 points across eight benchmarks and surpasses frontier models like Claude Opus 4.6 on AutomationBench, leading all compared 35B-A3B agent models.

Original post →

More from coding & agent

coding & agent channel →