Sai Agent Tops OSWorld 2.0 Benchmark at Lower Cost

TianbaoX · x · 2026-08-28

Simular's computer agent Sai achieved a 73% success rate on the OSWorld 2.0 benchmark, surpassing OpenAI's GPT-5.6 Sol (62.57%) and Anthropic's Opus 5 (70.57%). The benchmark consists of 108 complex tasks that take skilled humans over an hour to complete. Sai delivers this performance at roughly two-thirds the cost per task of its competitors. The key innovation is a neurosymbolic framework that pairs the exploratory power of neural networks with the logic of symbolic code, enhancing reliability and reducing costs.

Original post →

More from coding & agent

coding & agent channel →