Agent Made 119 Tool Calls on One Task, But the Benchmark Only Scored the Final Answer

patience__01 · reddit · 2026-09-05

In FinFIRST, Ling-3.0-flash-Fin averaged 23.08 reasoning rounds and 30.87 tool calls per task, hitting 119 calls on one task, with 122/123 valid outputs — yet the benchmark's GLM-5.1 judge scored only final answers, not reasoning or tool trajectories. The author argues this is a serious observability gap: answer quality says nothing about whether the path was inefficient, repetitive, or fragile, and agent evals must define what trajectory evidence to retain, score, and surface to operators.

Original post →

More from coding & agent

coding & agent channel →