Cua-S1-4B-0.2: First Multimodal Decision Model Trained with RLOO on Live Computer-Use Tasks, Apache-2.0
alexcovo_eth · x · 2026-09-24
Cua introduced Cua-S1-4B-0.2, the first multimodal decision model trained with RLOO on live computer-use tasks using task-completion rewards. Text and multimodal adapters are released under Apache-2.0; the open-source project (26k+ GitHub stars) ships cross-OS fleets, drivers, and benchmarks for training, evaluation, and data generation.
More from coding & agent
- Running a brand with only AI agents: the Notch experiment applied to nail polish ads — azed_ai · 2026-09-24
- "The replacement you trained just became your boss" — Codex gets StackOverflow plugin — cto_junior · 2026-09-24
- Splitting sandbox base layers with Nix: fewer images, more auditable agent environments — sloppenheimer · 2026-09-24
- Devin adds native Teams support and first-party Microsoft 365 integration — DevinAI · 2026-09-24
- Dev on AI sandbox tooling: audit chronus at syscall/eBPF layer, hide extraneous tools — sloppenheimer · 2026-09-24
- Notion Engineer: Subsidized Tokens Mean You Should Switch Agents Freely Without Losing Context — nbaschez · 2026-09-24