Eval tips for automated optimizers: use AUC over accuracy, sample 10x and check consensus to cut hallucinations
a_karvonen · x · 2026-09-28
Practical eval lessons for automated optimizers (AOs):
- Use AUC instead of accuracy: on a sycophancy eval, the AO scored at random chance on accuracy but much better with AUC, and became far less prompt-dependent.
- Sample multiple times and check consensus to mitigate hallucinations: AOs untrained to express uncertainty are confidently wrong; simply sampling 10 times and checking consensus yields a clean precision/recall curve.
Two cheap, immediately actionable eval improvements.
More from Research
- New blog surveys world models: definitions, SOTA, and eval axes — mervenoyann · 2026-09-29
- Researcher Flags Risk That Pedagogical RL's Student-Likelihood Optimization Filters Rare Reasoning — novasarc01 · 2026-09-29
- Goodfire grants geometric_intel lab funding for AI interpretability research — ninamiolane · 2026-09-29
- Bespoke Labs Launches AutoResearchExam Benchmark, Again, With a Demo Video — gregd_nlp · 2026-09-28
- Agentick benchmark accepted at NeurIPS: LLM vs RL agents on same tasks, no single winner — pcastr · 2026-09-28
- IROS 2026 has 1,933 papers — researcher curates 130-paper reading list on VLA and robot learning — GlenBerseth · 2026-09-28