First quantitative evidence: Claude and GPT now use GUIs as well as APIs
ysu_nlp · x · 2026-09-12
A shared benchmark post offers the first quantitative evidence that Claude and GPT models are rapidly progressing on computer-use agents: Fable 5.1 and GPT-6 Astra now operate GUIs as well as APIs, ending the "CUA tax," while Grok 4.6 degrades 71% on GUIs despite strong API performance.
More from Models
- DeepSeek V4.1 Flash shows massive kernel-engineering gains, hits 4th on KernelBench-CUDA — teortaxesTex · 2026-09-12
- AI cracks a Millennium Prize problem — proof of accelerating and alarming progress — pstAsiatech · 2026-09-12
- V4.1 Scores 11/70 on Terminal-Bench-Science, Strongest in Physical Sciences — teortaxesTex · 2026-09-12
- User pleads for boolean operators in Grok's chat search, which broadens instead of narrowing — chrisgrayson · 2026-09-12
- Google's upcoming model rumored to outperform GPT 5.6 Sol — imjustnewatai · 2026-09-12
- A real-world agent challenge: recreate a song's keyboard performance patch end to end — AdBest4099 · 2026-09-12