Meta's open-source AIRS-Bench featured in State of AI report, measuring AI R&D agents

j_foerst · x · 2026-10-08

AIRS-Bench, an AI R&D benchmark open-sourced by Meta AI, is featured in the 9th annual State of AI report alongside PostTrainBench. It measures agents' ability to execute end-to-end AI R&D across the full research lifecycle—idea generation, implementation, experiment analysis, and iterative refinement—and was used in the Muse Spark and Muse Spark 1.1 safety reports to assess risks of models automating AI R&D and outpacing governance. Quantifies both LLMs and harnesses in training AI models, matching the report's 'Gyms for AI' theme. Code, paper, and dataset are public.

Original post →

More from AGI Musings

AGI Musings channel →