HarnessOpt-Bench: New benchmark measures AI recursive self-improvement

shehio · reddit · 2026-08-28

Scale AI introduces HarnessOpt-Bench, a benchmark designed to measure Recursive Self-Improvement (RSI) by scoring an LLM on how much it improves another agent's harness. To prevent cheating (like the HF incident), the evaluation set is strictly isolated from the optimizer's sandbox. Testing 5 frontier models on 4 tasks reveals:

Model choice drives 1.8x more gain than harness choice.

Related event: Scale's HarnessOpt-Bench Measures Recursive AI Self-Improvement(2 posts)→

Original post →

More from Research

Research channel →