IBM's Tejas Kumar Warns of Four Common LLM Eval Pitfalls
IBM's Tejas Kumar delivered an hour-long evals workshop at AI Engineer, demonstrating four common LLM evaluation pitfalls and arguing that LLM judges need at least 80% agreement with humans before deployment.
2026-10-06 ~ 2026-10-07 · 2 related posts
- IBM's Tejas Kumar shows how substring matches and biased LLM judges break your evals — AI Engineer · 2026-10-06
- AI Engineer talk: LLM judge must match humans 80% before anything ships — TejasKumar_ · 2026-10-07