IBM's Tejas Kumar Warns of Four Common LLM Eval Pitfalls

IBM's Tejas Kumar delivered an hour-long evals workshop at AI Engineer, demonstrating four common LLM evaluation pitfalls and arguing that LLM judges need at least 80% agreement with humans before deployment.

2026-10-06 ~ 2026-10-07 · 2 related posts