HarnessOpt-Bench: New benchmark measures AI recursive self-improvement
shehio · reddit · 2026-08-28
Scale AI introduces HarnessOpt-Bench, a benchmark designed to measure Recursive Self-Improvement (RSI) by scoring an LLM on how much it improves another agent's harness. To prevent cheating (like the HF incident), the evaluation set is strictly isolated from the optimizer's sandbox. Testing 5 frontier models on 4 tasks reveals:
- Model > Harness: Swapping models yields significant gains (e.g., GPT from 3% to 49% headroom).
- No Home-Field Advantage: Models don't necessarily perform best in their native harnesses; the third-party OpenCode framework outperformed native ones in most pairs.
Model choice drives 1.8x more gain than harness choice.
Related event: Scale's HarnessOpt-Bench Measures Recursive AI Self-Improvement(2 posts)→
More from Research
- Michigan Robotics Rounds Up Its Papers and Workshops for IROS 2026 — doctorBobG · 2026-08-28
- Mouse Brain Connectome Cost Drops to $100M; Human at $1B — juanbenet · 2026-08-28
- Testing Muon Optimizer: Smoother Gradients and Stable Residual Maxima — stochasticchasm · 2026-08-28
- First open model adopts per-head orthogonalization following Kimi and GLM — stochasticchasm · 2026-08-28
- Integrity Bench: A New Benchmark to Measure Model Overconfidence — Acne_Discord · 2026-08-28
- Discussion on Classic Moonlight Scaling and Polar Express Orthogonalization — stochasticchasm · 2026-08-28