Meituan's Tulin team publishes comprehensive guide to Agent evaluation methodology and practice
美团技术团队 · wechat · 2026-08-06
Meituan's technical team published an in-depth blog on Agent evaluation, based on two years of practice by the Tulin Agent evaluation team. The article systematically covers the core purpose of Agent evaluation, differences from traditional model evaluation, methodology for building evaluation systems, and the evolution of observation and evaluation in the long-horizon Agent era.
Key points:
- Agent evaluation targets a complex system of "model + system + tools + process", covering result, process, efficiency, and risk layers.
- Observation is the foundation; Trace systems are needed to record full-path information.
- The core of an evaluation system is "bridging" business metrics and model capability metrics.
- Subjective evaluation requires alignment via "human-human consistency" and "human-machine consistency", breaking down vague metrics into binary rubrics.
- Evaluation is a practical science; iterate with Good/Bad cases rather than designing complex systems upfront.
- Expert knowledge can supplement model capabilities in vertical domains.
Case data: Digital station master achieved 99% human-machine consistency; Beam improved from 62% to 92% after adopting binary rubrics; fulfillment business expanded from 20+ to nearly 200 evaluation metrics.
The article also discusses changes brought by long-horizon agents (e.g., Claude Code, Lobster/Hermes), emphasizing that evaluation is evolving from a scoring action to an infrastructure capability.
More from coding & agent
- SKILL-KD: Contrastive Skill Distillation for Weaker LLM Agents — ZhejiangUniversity · 2026-08-06
- OneDayAgent: A Harness for Long-Horizon Autonomous Agents — zjunlp · 2026-08-06
- Developer Seeks Growth Advice for MCP-Based AI Meeting Scheduler — Early_Ad8608 · 2026-08-06
- AI Agents Replicate ICML Papers: Exposing Hidden Compute Costs and Flaws — MengdiWang10 · 2026-08-06
- Open-Source AI Skill: Generate Multi-Style ATS Resumes via Interview-Style Q&A — vista8 · 2026-08-06
- Core of Agent-to-Agent Trust: Verification Cost Dictates Architecture — anp2_protocol · 2026-08-06