LangChain releases Tuned Evaluators: Auto-score agents with 82% cost reduction

LangChain · x · 2026-08-19

LangChain released LangSmith Tuned Evaluators to automatically score agent behavior in production, starting with Perceived Error. Perceived Error is one of the clearest signals that an agent is providing a helpful experience. In LangChain's benchmark, their specialized model outperformed every frontier model tested and reduced evaluation costs by 82%. The tool offers a turnkey, managed service, removing the need for teams to write complex prompts, maintain LLM-as-a-judge models, or manage inference infrastructure.

Related event: LangChain launches Tuned Evaluators, cutting agent eval costs by 82%(4 posts)→

Original post →

More from coding & agent

coding & agent channel →