AgentLens Evaluates Coding Agent Trajectories

Andrey Podivilov · hf · 2026-07-10

AgentLens is a production-grade benchmark for coding agents that looks beyond task pass rates. It evaluates the agent's performance across its entire execution trajectory: instruction adherence, tool usage, self-verification, error correction, and user interaction.

The benchmark combines formal verification with LLM trajectory review and pairwise comparisons, ensuring every run yields a readable score explanation. The authors note that this methodology can be used not only for model rankings but also for diagnosing agent behaviors, comparing version differences, and catching product regressions in nightly evaluation pipelines. The project is open-source.

Original post →

More from coding & agent

coding & agent channel →