LLM judges are all finetunes of the same models, risking shared biases

deliprao · x · 2026-09-26

Deliprao argues that Jev from typesafeai is a pivotal product for judge models and AI design thinking, and the model style it advanced is here to stay. However, he points out that all Jev lookalikes appear to be finetunes of existing base models, meaning judges will make similar mistakes and reinforce the same biases in production. Switching to Jev itself won't fix this either — hence the need to build competent judges afresh to ensure true diversity.

Related event: Delip Rao: LLM Judges Are Mostly Fine-tunes, Correlated Errors Limit Ensemble Gains(5 posts)→

Original post →

More from Models

Models channel →