AutoBenchmark: benchmark difficulty transfers to held-out models like Nemotron and Claude Opus-5

jaseweston · x · 2026-09-30

Part 4/5 of the AutoBenchmark thread: benchmarks built with Muse Spark and Glimmer as in-loop solvers transfer in difficulty to held-out models, Nemotron and Claude Opus-5. The same trends hold — human aid helps, and without it benchmarks saturate.

Related event: Meta AI introduces AutoBenchmark for automated benchmark creation(3 posts)→

Original post →

More from Research

Research channel →