HKU Introduces OS-Shepherd: Low-Cost, Reliable Reward Models for CUAs
hkunlp · hf · 2026-08-07
The HKUNLP team at HKU released OSReward, a new benchmark addressing the evaluation challenges of Computer-Using Agents (CUAs).
- Background & Pain Points: Verifying CUA trajectories heavily relies on Vision-Language Models (VLMs) as judges. However, current VLMs exhibit a systematic leniency bias, often mislabeling failed runs as successes. Reliable commercial models are too expensive, while affordable open-source ones lag significantly behind.
- Benchmark & Dataset: The team built OSReward with real-world trajectories, a challenge set OSReward-Hard, and released OS-Shepherd-100K, a corpus of 100K reasoning-annotated trajectory judgments.
- Models & Results: They trained open reward models OS-Shepherd (9B and 35B), which provide stable and reliable reward signals at a 30-60% lower cost than frontier commercial judges.
More from coding & agent
- Testing Claude Opus with Unity CLI: A Major Breakthrough in 3D Spatial Understanding for Gamedev — chongdashu · 2026-08-07
- Are MCP Servers Becoming Architectural Dependencies? Devs Worry About Portability — dancepeop · 2026-08-07
- Agent Harness Bloat is Real: Minimal Context Cuts Costs and Runs — zainhas · 2026-08-07
- Indie Dev Clones Influencer Voice with VoxCPM, Achieving 3s End-to-End Latency — 面壁智能 · 2026-08-07
- Local Models Output Gibberish in Agent Mode: Why Ability Boundaries Matter — Marblapas · 2026-08-07
- DeepSeek Cuts Agentic Loop Costs 100x Without Quality Loss — bindureddy · 2026-08-07