ClawProBench: Trace-Aware Agent Evaluation with Runtime Coverage and Frozen Tasks

YuanHang Xiao · hf · 2026-08-26

A new benchmark, ClawProBench, has been released on Hugging Face for evaluating AI agents via execution traces.

Core Features:

This benchmark aims to provide a more accurate measure of agent configurations in real-world workflows.

Original post →

More from coding & agent

coding & agent channel →