A controlled study finds no universal best harness for automated discovery
LChoshen · x · 2026-07-23
Automated discovery does not have a universally best harness
A study shared in the thread reports that no single harness was reliably best for automated discovery across model–problem pairs.
Key findings:
- Harness choice behaves like a model/problem-specific hyperparameter, not a universal recipe.
- Simpler harnesses often matched or beat more complex systems, including OpenEvolve variants.
- The best harness changed depending on the model–problem pair.
- Because this setting is high variance, the authors used repeated budget-matched runs, strong baselines, and statistical hypothesis testing to compare methods robustly.
The practical takeaway is that online harness selection may matter more than adding complexity to the discovery system itself.
Related event: Study Finds No Universal Optimal Harness for Automated Discovery(2 posts)→
More from Research
- 3D ResNet Paper Crosses 3,000 Citations Eight Years After CVPR 2018 — HirokatuKataoka · 2026-09-11
- Sample selection and ordering matter a lot in LLM training: DataFlex makes data scheduling dynamic — Puzzleheaded_Box2842 · 2026-09-11
- Jeff Heaton's Intro to the Math of Neural Networks eBook Is Free to Download — blaizedsouza · 2026-09-11
- Mathematician Daniel Litt Launches Problem Repo to Track Human vs AI Progress: 15 Problems, 1 Solved — littmath · 2026-09-11
- Open ECDSA.fail challenge uses AI agents to shrink Shor's-algorithm quantum circuits for Bitcoin keys — StefanoGogioso · 2026-09-11
- Alex Townsend posts 200 open problems in numerical linear algebra for humans and AI agents — IgorCarron · 2026-09-11