Meta AI's AutoBenchmark auto-creates benchmarks, finds human feedback gives big wins over agents alone
jaseweston · x · 2026-09-30
Meta AI researcher Jason Weston introduced AutoBenchmark: an agentic system that creates benchmarks automatically — specifically benchmarks that evaluate autoresearch agents themselves — closing a full recursive-improvement loop, while studying the role of humans in the loop.
Key findings: human-agent collaboration substantially beats agents alone, with fine-grained human feedback during the ideation stage crucial; autoresearch benchmarks for AI research can be built with this recipe; and benchmark creation works best with two feedback sources — benchmark solvers plus external verifiers (human + AI).
More from Research
- Missing API for general real-time LLM agents: AsyncLLM preprint sparks interface debate — phill1992 · 2026-09-30
- Cohere Labs to host IOL-AI 2026 wrap-up on why linguistic reasoning still stumps models — Cohere_Labs · 2026-09-30
- UK's Zenithon raises $10M to build world models for extreme physics: rockets, fusion and fabs — roydanroy · 2026-09-30
- Tenstorrent opens bio-model training: OpenFold3 on Blackhole Galaxy nears DGX H200 at quarter the cost — MoAlQuraishi · 2026-09-30
- MIT's Ataraxo AI beats top Stratego players with self-play and decision-time planning — nordicinst · 2026-09-30
- New Research: AI as Tutor Beats AI as Substitute — and No AI — CackleRooster · 2026-09-30