Cua Releases Cua-S1-4B, First Multimodal Decision Model RL-Trained on Live Computer-Use Tasks

multimodalart · x · 2026-09-24

trycua introduced Cua-S1-4B-0.2, billed as the first multimodal decision model trained with RLOO on live computer-use tasks using task-completion rewards. Text and multimodal adapters are Apache-2.0 licensed, and a Hugging Face Space lets you try the model in your browser and see how it scores candidate actions.

Related event: Cua Releases Cua-S1-4B, First RL-Trained Computer-Use Decision Model(2 posts)→

Original post →

More from coding & agent

coding & agent channel →