Cua Releases Cua-S1-4B, First RL-Trained Computer-Use Decision Model

trycua released Cua-S1-4B-0.2, claimed to be the first multimodal decision model trained with RLOO reinforcement learning on real computer-use tasks, using task completion as the reward.

2026-09-24 ~ 2026-09-24 · 2 related posts