Simular's Agent Sai Tops OSWorld 2.0, Beating GPT-5.6 at Two-Thirds the Cost

xwang_lk · x · 2026-08-28

Simular's computer agent Sai achieved a 73% success rate on OSWorld 2.0, ahead of GPT-5.6 Sol (62.57%, per OpenAI) and Opus 5 (70.57%, per Anthropic), at roughly two-thirds the cost.

OSWorld 2.0, released in June by HKU's XLANG Lab, has 108 long-horizon tasks taking skilled humans over an hour each: cross-source reasoning (receipts scattered across email and expense reports), responding to dynamic changes, following tutorials precisely, and troubleshooting contradictory data.

Sai operates desktop apps and webpages, calls APIs, and writes code like a human — orchestrating a mixture of frontier and specialist models with dedicated perception/action interfaces, running on Simular's remote fleet or personal machines.

Related event: Simular's Sai Agent Tops OSWorld 2.0 at Two-Thirds the Cost(4 posts)→

Original post →

More from coding & agent

coding & agent channel →