Meta AI's AutoBenchmark auto-creates benchmarks, finds human feedback gives big wins over agents alone

jaseweston · x · 2026-09-30

Meta AI researcher Jason Weston introduced AutoBenchmark: an agentic system that creates benchmarks automatically — specifically benchmarks that evaluate autoresearch agents themselves — closing a full recursive-improvement loop, while studying the role of humans in the loop.

Key findings: human-agent collaboration substantially beats agents alone, with fine-grained human feedback during the ideation stage crucial; autoresearch benchmarks for AI research can be built with this recipe; and benchmark creation works best with two feedback sources — benchmark solvers plus external verifiers (human + AI).

Related event: Meta AI's AutoBenchmark auto-creates benchmarks, finds human feedback gives big wins over agents alone(3 posts)→

Original post →

More from Research

Research channel →