NVIDIA's OSWorld-Pro adds 2,800+ subgoal process evaluation — Claude Opus 5 drops to 75.7%
nvidia · hf · 2026-10-06
- Existing computer-use agent (CUA) benchmarks like OSWorld assess only final deliverables with functional verifiers, hiding how and why agents fail.
- OSWorld-Pro offers 300+ tasks with 2,800+ subgoals grounded in 67,000+ human annotations, using human-aligned LLM-Judges to score fulfillment of sequentially dependent subgoals for procedural evaluation.
- Key results: it's challenging even for SOTA LLMs — Claude Opus 5 scores 83.4% on OSWorld but only 75.7% on OSWorld-Pro.
- Analysis reveals process-focused failure modes (subgoal-irrelevant actions, click-based mistakes), offering insights to improve CUA performance and efficiency.
More from Research
- Cornell Tech prof presents two pretraining-behavior papers at COLM, recruiting PhDs and postdocs — pratyushmaini · 2026-10-06
- Non-Expert Runs AI-Driven Interpretability Study Claiming Linear Representation of Directive Force in LLMs — LoudYogurtcloset7856 · 2026-10-06
- AC2 lets LLM RL train faster than GRPO by trusting critics on partial rollouts — stanfordnlp · 2026-10-06
- Stanford CS224N Winter 2026 posts free slides and assignments online — stanfordnlp · 2026-10-06
- Only right answers still teach chemistry: popular chemistry benchmark has shortcuts — tak3sh8 · 2026-10-06
- UT Austin math chair: OpenAI appears set to release ~400 AI-generated proofs at once — 141_1337 · 2026-10-06