Agents fail to reconsider strategy during post-training execution

omarsar0 · x · 2026-08-20

A paper tests if agents can effectively post-train other agents. It finds that agents lock in their training strategy at the first step and spend the rest of the budget on local adjustments. Experience-driven scaffolds improved execution significantly (+12.6 on GSM8K, +40.8 on HumanEval) but kept the strategy frozen. Human guidance redirected the opening choice but led back to local loops. Extra inference compute helped only on easy tasks. The key missing piece is the ability to reconsider strategy during execution.

Original post →

More from coding & agent

coding & agent channel →