IBM's Tejas Kumar shows how substring matches and biased LLM judges break your evals
AI Engineer · youtube · 2026-10-06
AI Engineer published a 1h workshop by IBM's Tejas Kumar on building evals that survive production, via live demos:
- Opening failure: a support answer passes a test merely for containing "cannot", while actually approving a return against store policy — where assertions and fuzzy matching break down
- Escalation path: builds an LLM judge, then a dataset of scenarios, agent responses, and human verdicts; the judge turns out to favor answers from its own model family (self-preference)
- Four judge failure modes: position bias, sycophancy, self-preference, verbosity — the instrument itself needs calibration against human verdicts
- Engineering practices: an 80% agreement threshold as a CI gate; keep datasets current with fresh traffic and human review; separate pre-deployment checks from runtime harness protections; start with deterministic checks before paying for model judgments
- OpenRAG demo: retrieves the latest uploaded refund policy instead of static prompt rules
Takeaway: a green test isn't proof the system works — keep reviewing disagreements, supplying context, and swapping models when needed.
Related event: IBM's Tejas Kumar Warns of Four Common LLM Eval Pitfalls(2 posts)→
More from coding & agent
- Open-Source Project With 2,000+ Stars Adds 9 AI Animation Explainer Styles and Full Workflow — dotey · 2026-10-07
- ChatGPT Launches Meetings Plugin That Auto-Notes Calls and Feeds Context Into Codex — craigsdennis · 2026-10-07
- Luma events via MCP can't add custom form questions, forcing users back to the browser — aronkor · 2026-10-07
- "Your agents want to work in a monorepo" — and the Bazel joke that followed — vboykis · 2026-10-07
- MCP Events UX is painful for consumer apps, developer reports six-step activation flow — haltakov · 2026-10-07
- Tip: force GPT-Live to say exact phrases via a developer message in session.start's input — pbbakkum · 2026-10-07