Stanford: 50 Examples Suffice to Evaluate Large Audio Models, HUMANS Benchmark Open-Sourced
stanfordnlp · x · 2026-09-16
Stanford NLP researchers published "Putting HUMANS first: Efficient LAM Evaluation with Human Preference Alignment," on efficient evaluation of large audio models (LAMs).
- Analyzing 10 subset selection methods, 18 audio models, and 40 tasks, they show that subsets of just 50 examples (0.3% of data) achieve over 0.93 Pearson correlation with full benchmark scores.
- From 776 human preference ratings on realistic voice assistant conversations, both subsets and the full benchmark only reach 0.85 correlation with user satisfaction.
- Regression models trained on curated subsets hit 0.98 correlation with human preference, beating random subsets and the full benchmark — quality over quantity.
They open-source these regression-weighted subsets as the HUMANS benchmark.
More from Research
- Google's Retrieve-for-Train replaces heavy autoregressive inference with a lightweight RL-trained diffusion model — gaganghotra_ · 2026-09-16
- OpenAI's Neon: 1,300 H200s and a materials lab loop topped GPT-6 Astra on analysis benchmark — daniel_mac8 · 2026-09-16
- "Potemkin Understanding": LLMs ace definitions but collapse on spotting real examples — anselm · 2026-09-16
- ChatGPT Co-Inventor Launches Jev After 2 Years in Stealth, Claiming 20-200x Speed and 40-400x Cost Gains — sedielem · 2026-09-16
- Periodic Neon beats GPT-6 Astra and Claude Fable 5.1 on XRD analysis at lower cost — zainhas · 2026-09-16
- Latent Spacecraft Project Links Brain Language Mechanisms to GAN Latent Spaces via Joyce — begusgasper · 2026-09-16