Evaluating Agents Without Right Answers: Similarweb's Playbook

LangChain · x · 2026-07-30

Similarweb detailed their engineering experience using LangSmith to evaluate long-form Agent research reports.

The core challenge is that unlike traditional software with clear pass/fail tests, agentic systems can take different paths and produce varying valid outputs for the same input. To tackle this, they adopted a multi-dimensional evaluation strategy:

The author emphasizes treating scores as signals, not absolute answers, and calibrating rubric weights carefully before use to avoid misjudging system improvements as regressions.

Original post →

More from coding & agent

coding & agent channel →