Harness-only changes lift deepagents-cli from 52.8% to 66.5% on Terminal-Bench 2.0
Gauri_the_great · x · 2026-09-07
Agent harnesses and infrastructure are an active engineering field. LangChain kept GPT-5.2-Codex fixed and lifted deepagents-cli on Terminal-Bench 2.0 from 52.8% to 66.5% using harness changes only: injecting environment context, enforcing a build–verify–fix loop, adding loop/timeout guards, and allocating reasoning budget by phase. The author argues that for agents running for hours, checkpointing, sandboxing, retries, tool permissions, tracing, and verification determine success — model intelligence is only one part. Ends with a joke that Fable 5.1 one-shots the whole task anyway.
More from coding & agent
- Astra sparks AI-CAD wave: creators rebuild landmarks and engines across Blender, SOLIDWORKS, Fusion — burhop · 2026-09-07
- OpenAI internal data: over 80% of successful 32-hour agent tasks still needed human intervention — DataLearnerAI · 2026-09-07
- threepointone: UX is about depth, agent experience is about breadth — threepointone · 2026-09-07
- dhh migrated a whole site to Astro with Muse Code for just 20 cents — alexandr_wang · 2026-09-07
- Microsoft dev blog: building AX evals that actually work, when scalar scores mislead — lee_stott · 2026-09-07
- UMass AutoIndex turns chunking into code an LLM writes, with hypothesis-level validation lift — CShorten30 · 2026-09-07