Stanford paper: fix looping agents via harness edits, bad plans need weight training
rohanpaul_ai · x · 2026-10-11
A Stanford paper (arXiv 2610.11655) proposes a practical decision rule for improving long-horizon LLM agents: label why runs fail first, then pick the right lever.
- Failures split into process failures (loops, blocked calls, exhausted step budgets) and content failures (a delivered but poor plan).
- On the DeepPlanning benchmark, a self-evolving harness loop lifted Qwen3.5-4B's held-out score from 0.16 to 0.30 and 9B from 0.32 to 0.44; for 4B, plan delivery rose from 55% to 90%. But harness edits never reduced the share of poor plans.
- LoRA adapters trained on evolved-harness trajectories internalized the gains: content failures on 9B dropped from a quarter of trajectories to one in twenty; on 4B the adapter stacked with the harness to more than double held-out scores.
More from coding & agent
- TikTok turns on vibe-coded Photoshop alternative; ex-Adobe dev says no one knows what's in Adobe's code either — jasonkneen · 2026-10-12
- Codex roadmap leak shows agent message board among 59 features in development — gajesh · 2026-10-12
- Claude Code creator: Opus 5.5 does a month of work in a day — and your old prompts now backfire — VeryWellVersed · 2026-10-12
- Solo dev ships Scape: Claude Code executes, Codex reviews in adversarial loop — croovies · 2026-10-12
- banteg builds a DSL to describe all 50 original game quests — banteg · 2026-10-12
- Bittensor SN33 claims skills turn small models into frontier-level performers — markjeffrey · 2026-10-12