Missing Benchmark for Agent Runtimes, Not Just Models, Reddit Thread Argues

Balance- · reddit · 2026-09-11

A Reddit poster notes that benchmarks like SWE-bench, Terminal-Bench, OSWorld, and BrowseComp all evaluate models, but no apples-to-apples benchmark exists for agent runtimes like OpenAI Agents, Anthropic's agent stack, AWS AgentCore, Google's agent platform, or LangGraph.

He proposes measuring task success rate, cost per successful task, wall-clock time, tool calls/retries, reliability over long-running tasks, and human interventions required. The most interesting experiment would control both sides: same model with different harnesses, and same harness with different models — since the industry increasingly evaluates "model + harness" systems while benchmarks still treat the model as the unit of comparison.

Original post →

More from coding & agent

coding & agent channel →