Swapping harness lifts GPT-5.6 repo migration from 6.5% to 31%, paper finds
omarsar0 · x · 2026-10-09
A new paper argues harness engineering is underexplored: with the same model and effort, replacing Codex with the HERMES harness lifts GPT-5.6 Sol from 6.5% to 31.0% on whole-repository migration.
- Design: HERMES gives each repository component a resident LLM that knows its own code and dependencies; a dependency-aware step decides which components to activate, and a diagnosis step maps test failures back to the components needing changes.
- Results: Beats matched baselines by 12.4 points on average across four software engineering benchmarks.
- Cost: With strong activation/diagnosis models, Qwen3-8B components come within 4.5 points of an all-GPT-5.6 setup while cutting Terminal-Bench 4.0 inference cost by 26.2%.
Takeaway: "build your own harness" — the scaffold can matter as much as the model.
More from coding & agent
- app-store-screenshots hits 7.2k GitHub stars: AI agent skill scaffolds store-ready screenshot editor — tom_doerr · 2026-10-09
- LLM-as-a-Verifier: Weaker Model Verifies Stronger One, Hits 69.2% SOTA on Terminal-Bench 4 — Azaliamirh · 2026-10-09
- TestSprite Season 4 developer contest offers $4,000 prize pool for Claude Code and Codex users — JaynitMakwana · 2026-10-09
- Salesforce's SRD distills hindsight into foresight, lifting 2B agent success from 0% to 60.6% — Salesforce · 2026-10-09
- Prompt Tuning Is Forgotten Lore — Are We Massively Underusing Finetuned Tokens? — cephaloform · 2026-10-09
- System architect shares WhatsApp voice-note AI assistant setup — No_Kangaroo_4454 · 2026-10-09