OSS harness boosts OS World 2.0 score ~20%, jumping Sol/Max from 4th to 1st
demeyer1 · reddit · 2026-09-15
A developer built an MIT-licensed open-source agent harness optimized for long-duration, high-complexity knowledge work, and quantified how much a harness alone can move a hard benchmark.
Running the full formal submission procedure on OS World 2.0, they compared Sol/Max standalone versus Sol/Max with the harness, publishing a full audit package to verify no cheating. The harness lifted Sol/Max from 4th place to 1st, ahead of Opus 5 — roughly "skipping a full generation of frontier models," per the author. Everything is open source with no benchmark-specific optimizations, and the author is soliciting other OSS harnesses with similar major-benchmark impacts.
More from coding & agent
- Musecases hits 101 real agent use cases, from AT&T bill negotiation to IKEA returns — thisiskp_ · 2026-09-15
- Delimit dev argues agent PR delegation should be scoped by destination repos, not diff correctness — delimitdev · 2026-09-15
- Playground launches a shared canvas for PMs, founders and agents, pitched by a16z — _AustinCalvert_ · 2026-09-15
- Unprompted 2026 opens virtual tickets for two days of AI x Cybersecurity research — moyix · 2026-09-15
- Cline ships native desktop app beta: coding agent without an editor — aigclink · 2026-09-15
- AI thriller 'All Aboard': 9 generated blocks, color-coded characters on Seedance 2.5 — tsi_org · 2026-09-15