Sai Agent Tops OSWorld 2.0 Benchmark at Lower Cost
TianbaoX · x · 2026-08-28
Simular's computer agent Sai achieved a 73% success rate on the OSWorld 2.0 benchmark, surpassing OpenAI's GPT-5.6 Sol (62.57%) and Anthropic's Opus 5 (70.57%). The benchmark consists of 108 complex tasks that take skilled humans over an hour to complete. Sai delivers this performance at roughly two-thirds the cost per task of its competitors. The key innovation is a neurosymbolic framework that pairs the exploratory power of neural networks with the logic of symbolic code, enhancing reliability and reducing costs.
More from coding & agent
- Take: context engineering starts at product UX design time, not at build time — andreisavu · 2026-08-28
- Engineering Discussion: Optimizing Context Compaction in DeepSeek Harness — shady101852 · 2026-08-28
- Microsoft Foundry makes Agent Hosting a first-class .NET primitive — adnan_hashmi · 2026-08-28
- Dev Seeks Complex Web App Specs to Stress Test Autonomous Coding Agent — ZenenoDev · 2026-08-28
- Florence launched: AI-ready design system for agents — Scobleizer · 2026-08-28
- AI agents explode issue backlogs, demanding new workflow — rseroter · 2026-08-28