UCSC's Prediction-Powered Smoothing Makes Disaggregated AI Evaluation More Precise
UCSantaCruz · hf · 2026-09-22
A UC Santa Cruz team tackles imprecise disaggregated AI evaluation — where performance varies across benchmark task types or conversation types but labeled samples per domain are scarce — with a workflow built on small area estimation.
- Estimation: prediction-powered smoothing (PP-S) fits a Bayesian model to each domain's prediction-powered estimate; an extension, PP-TS, borrows strength across a reporting taxonomy.
- Validation: a new approximately unbiased design-based cross-validation score selects between direct and smoothed estimators.
- Experiments: validated on two settings with fully observed outcomes — a curated benchmark with verifiable grading and deployed agent traffic graded by humans. The proposed estimators beat direct estimators in point and interval estimation with near-nominal coverage.
- Takeaway: at equal sampling budget, the score selects as well as an independent validation sample and estimates the chosen estimator's error far more accurately.
More from Research
- Science Paper: Quantum-Optical Spin Glass Delivers High-Cavity Hopfield Associative Memory — SuryaGanguli · 2026-09-22
- Room-temperature superconductor claims resurface: wrinkled graphite nanoflakes trap magnetic flux above 390K — skdh · 2026-09-22
- Math isn't as open as claimed: NSA kept differential cryptanalysis secret since 1974 — lpachter · 2026-09-22
- New Blog Explores Group Theory Meets Deep Learning with Interactive Sphere Visuals — CatAstro_Piyush · 2026-09-22
- Qonto open-sources QontoFAQ, a retrieval benchmark closer to real product Q&A — espadrine · 2026-09-22
- Podcast: Epoch AI researcher on RSI, robotics, China model gap, and open vs closed safety — natolambert · 2026-09-22