Cua-S1-4B-0.2: First Multimodal Decision Model Trained with RLOO on Live Computer-Use Tasks, Apache-2.0

alexcovo_eth · x · 2026-09-24

Cua introduced Cua-S1-4B-0.2, the first multimodal decision model trained with RLOO on live computer-use tasks using task-completion rewards. Text and multimodal adapters are released under Apache-2.0; the open-source project (26k+ GitHub stars) ships cross-OS fleets, drivers, and benchmarks for training, evaluation, and data generation.

Original post →

More from coding & agent

coding & agent channel →