Delip Rao: finetuned judge models like Jev won't deliver true judge diversity
deliprao · x · 2026-09-26
AI researcher Delip Rao called typesafeai's Jev a pivotal product and design direction, but cautioned that all Jev-like models appear to be finetunes of existing LLMs. That means the resulting "judges" make similar mistakes and reinforce the same biases in production — so switching to this version of Jev doesn't help. He argues we need to build such models from scratch to get true diversity of competent judges.
Related event: AI Judge Models Are Mostly Fine-Tunes, Raising Bias Concerns(2 posts)→
More from Models
- Ethan Mollick: 'Keep prompts short' is bad advice, and minimizing token cost confuses inputs with outputs — emollick · 2026-09-26
- GPT-6 Luna uses fewer reasoning tokens than 5.6 on ARC-AGI-2, hard tasks stymie both — mhmazur · 2026-09-26
- Heavy user: fast, cheap Claude Opus 5.5 now takes all my serious work — brandon_galang · 2026-09-26
- Reddit users say Claude Opus 5.5 is 'stupidly good' at game creation — Adorable_Being_763 · 2026-09-26
- Model excels at Sokoban-style puzzles, sparking questions about training data contamination — lukaszkaiser · 2026-09-26
- GPT-6 Astra's First Draft Fooled Every's CEO: Big Writing Upgrade, But Some Bad Habits — every · 2026-09-26