Experiments show smarter models and higher effort write better LLM-judge evals
danshipper · x · 2026-10-08
An ongoing experiment on automated evals asks whether more compute means better checks from LLM judges — and the answer is 'yes, kind of.' Smarter models wrote better evals, and raising reasoning effort notably improved eval quality for small models like Luna.
More from Research
- EMNLP 2026 paper RECAP trains reasoning models to recover from unsafe trajectories — pinyuchenTW · 2026-10-08
- SaTML 2027 to host AdvML Frontiers workshop on human-centered trustworthy machine learning — pinyuchenTW · 2026-10-08
- OpenAI's new result proves 2005 edit-distance embedding optimal; researcher distills proof to 2.5 pages with AI help — thegautamkamath · 2026-10-08
- Tencent's WorkForge scales verifiable training environments for long-horizon work agents — teortaxesTex · 2026-10-08
- Masked Geometric Encoder boosts 3D foundation models via frame dropping and self-distillation — zhenjun_zhao · 2026-10-08
- DensiTok: flow-matching token densification lets frozen feed-forward 3DGS see unseen views — zhenjun_zhao · 2026-10-08