Cutting Evaluation Costs via Black-Box Predictions
BlackHC · x · 2026-07-10
This post introduces a new approach to active testing in model evaluation: estimating model risks with fewer labels while leveraging cheap predictions from strong black-box models as an additional resource.
The authors propose the PPAT method to mitigate the high annotation costs associated with traditional evaluation.
More from Research
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22
- Ai2’s Asta adds one-click handoff and self-checking deep paper search — allen_ai · 2026-07-22
- DepthART pushes monocular depth to tiny models at 1000 FPS on RTX A6000 — kwangmoo_yi · 2026-07-22
- Meta says SAM 3 and DINOv3 cut 3D volume labeling from a month to 15 minutes — AIatMeta · 2026-07-22
- Project CETI gets a Jeopardy! shout-out with a SETI-style whale clue — begusgasper · 2026-07-22