WebStep: A New Benchmark for Process-Level Evaluation of Web Agents
algo_diver · x · 2026-08-13
Accepted at COLM 2026, WebStep introduces a process-level evaluation method for web agents. Existing benchmarks typically focus only on terminal success rates, but two agents can fail the same task for entirely different reasons.
WebStep introduces an MDP observer that automatically translates low-level GUI actions (like click coordinates) into semantic actions and states (e.g., ViewRepo). The benchmark consists of 1,800 task instances across 10 self-hosted websites.
Using the recorded traces, WebStep computes process metrics such as exploration reach, skill invocation, and execution efficiency. It pinpoints exactly where a failing trajectory went wrong, revealing behavioral differences often hidden by similar overall success rates.
More from coding & agent
- KOF Nano Banana MCP Server: Enables Batch Image Generation with Gemini via YAML — modelcontextprotocol · 2026-08-13
- Fast.io Launches MCP Toolkit: 251 File Collaboration and RAG Tools for AI Agents — modelcontextprotocol · 2026-08-13
- Multi-Model Workflow: Assigning Different LLMs for Coding, Architecture, and Review — jiayuan_jy · 2026-08-13
- A Thumbs-Up Emoji Broke My AI Sales Agent: Don't Let 'Safe' Fallbacks Be the Costliest Error — Sudden-Theme7554 · 2026-08-13
- Grok Voice Agents Can Now Autonomously Handle End-to-End Customer Support — XFreeze · 2026-08-13
- AI-Generated 'Load-Bearing' Code Surges in Intercom's Rails Monolith — teropa · 2026-08-13