Debating agent evals: traces over UI

zeeg argues that evaluating agents should rely on full session traces and tool calls via vitest-evals rather than UI, positioning OpenEval as a more isolated prompt-testing variant compared to his hybrid testing approach.

2026-09-12 ~ 2026-09-12 · 2 related posts