IBM's Tejas Kumar shows how substring matches and biased LLM judges break your evals

AI Engineer · youtube · 2026-10-06

AI Engineer published a 1h workshop by IBM's Tejas Kumar on building evals that survive production, via live demos:

Takeaway: a green test isn't proof the system works — keep reviewing disagreements, supplying context, and swapping models when needed.

Related event: IBM's Tejas Kumar Warns of Four Common LLM Eval Pitfalls(2 posts)→

Original post →

More from coding & agent

coding & agent channel →