AutoBenchmark wraps up: human direction helps when autoresearch loops stall, plus Rebuttal Bench
jaseweston · x · 2026-09-30
Final part of the AutoBenchmark thread (with full Meta AI blog post): when the autoresearch loop stalls, human direction at that point helps. The team also introduces Rebuttal Bench (assessment) — testing whether a paper rebuttal actually resolves reviewer weaknesses. The blog details the method: autoresearch agents iteratively create benchmarks with feedback from solvers and external verifiers (human + AI), studying the scope and necessity of human contribution in open-ended tasks.
Related event: Meta AI introduces AutoBenchmark for automated benchmark creation(3 posts)→
More from Research
- Missing API for general real-time LLM agents: AsyncLLM preprint sparks interface debate — phill1992 · 2026-09-30
- Cohere Labs to host IOL-AI 2026 wrap-up on why linguistic reasoning still stumps models — Cohere_Labs · 2026-09-30
- UK's Zenithon raises $10M to build world models for extreme physics: rockets, fusion and fabs — roydanroy · 2026-09-30
- Tenstorrent opens bio-model training: OpenFold3 on Blackhole Galaxy nears DGX H200 at quarter the cost — MoAlQuraishi · 2026-09-30
- MIT's Ataraxo AI beats top Stratego players with self-play and decision-time planning — nordicinst · 2026-09-30
- New Research: AI as Tutor Beats AI as Substitute — and No AI — CackleRooster · 2026-09-30