Cua offers scaled computer-use agent fleets, but top frontier agent clears just 6 of 25 KiCad tasks
lucasmeijer · x · 2026-09-11
Lucas Meijer asked what to know before wiring computer/browser use into his agents, pointing to Cua: a platform for running computer-use agent training, evals, and data generation at scale across Linux, Windows, macOS, and Android VMs, with snapshot forking, failure reproduction, and fleet pools that scale to zero. Its Cua-Bench shows the best frontier agent clears only 6 of 25 expert KiCad tasks — a sobering signal on real-world computer-use reliability.
More from coding & agent
- Devs still debug agent runs for hours while demos promise autonomous research — DominiqueCAPaul · 2026-09-11
- AI Writes Entire 3D Game Engine Overnight in Bend2, Hitting 120 FPS — rickasaurus · 2026-09-11
- Vesence launches agent-native browser desktop with Office file editing and approval gates — garrytan · 2026-09-11
- GPT-6 Given 2 Hours to Build a Viral Site Makes a Draw-Your-Horse Racing Game — yungcontent · 2026-09-11
- PARSER: parallel chunk subagents with an RL-trained lead agent for long-context QA — omarsar0 · 2026-09-11
- Everyone ships more than ever, but revenue isn't keeping up, says Clairevo CEO — lennysan · 2026-09-11