AutoBenchmark wraps up: human direction helps when autoresearch loops stall, plus Rebuttal Bench

jaseweston · x · 2026-09-30

Final part of the AutoBenchmark thread (with full Meta AI blog post): when the autoresearch loop stalls, human direction at that point helps. The team also introduces Rebuttal Bench (assessment) — testing whether a paper rebuttal actually resolves reviewer weaknesses. The blog details the method: autoresearch agents iteratively create benchmarks with feedback from solvers and external verifiers (human + AI), studying the scope and necessity of human contribution in open-ended tasks.

Related event: Meta AI introduces AutoBenchmark for automated benchmark creation(3 posts)→

Original post →

More from Research

Research channel →