Delip Rao: finetuned judge models like Jev won't deliver true judge diversity

deliprao · x · 2026-09-26

AI researcher Delip Rao called typesafeai's Jev a pivotal product and design direction, but cautioned that all Jev-like models appear to be finetunes of existing LLMs. That means the resulting "judges" make similar mistakes and reinforce the same biases in production — so switching to this version of Jev doesn't help. He argues we need to build such models from scratch to get true diversity of competent judges.

Related event: AI Judge Models Are Mostly Fine-Tunes, Raising Bias Concerns(2 posts)→

Original post →

More from Models

Models channel →