Cua Releases Cua-S1-4B, First RL-Trained Computer-Use Decision Model
trycua released Cua-S1-4B-0.2, claimed to be the first multimodal decision model trained with RLOO reinforcement learning on real computer-use tasks, using task completion as the reward.
2026-09-24 ~ 2026-09-24 · 2 related posts
- Cua-S1-4B-0.2: First Multimodal Decision Model Trained with RLOO on Live Computer-Use Tasks, Apache-2.0 — alexcovo_eth · 2026-09-24
- Cua Releases Cua-S1-4B, First Multimodal Decision Model RL-Trained on Live Computer-Use Tasks — multimodalart · 2026-09-24