Agent Made 119 Tool Calls on One Task, But the Benchmark Only Scored the Final Answer
patience__01 · reddit · 2026-09-05
In FinFIRST, Ling-3.0-flash-Fin averaged 23.08 reasoning rounds and 30.87 tool calls per task, hitting 119 calls on one task, with 122/123 valid outputs — yet the benchmark's GLM-5.1 judge scored only final answers, not reasoning or tool trajectories. The author argues this is a serious observability gap: answer quality says nothing about whether the path was inefficient, repetitive, or fragile, and agent evals must define what trajectory evidence to retain, score, and surface to operators.
More from coding & agent
- Devs call to abolish context compaction: agent engineering debate heats up — sull · 2026-09-05
- GPT-6 Astra arrives in GitHub Copilot, built for long-horizon autonomous coding — OpenAIDevs · 2026-09-05
- Hierarchical RAG plus persistent scratchpad beats FIFO windows for long-horizon agents — IgorCarron · 2026-09-05
- Astra-built site lays out GPT-6's pitch: research, modeling, coding, cross-app agents — reach_vb · 2026-09-05
- Dev Uses 5.6-luna-low Model to Simulate Dumb Users Before Shipping — generativist · 2026-09-05
- Everyone Asks Astra For Games, This Dev Asked It To Deflake Tests — cramforce · 2026-09-05