Self-verify-and-repair loop for coding agents lifts Terminal-Bench score by 15 points

Sorosu · reddit · 2026-08-19

A Reddit user shares the open-source Autoprompt skill, which wraps coding agents in a full plan → implement → test → review → repair loop so the agent autonomously verifies and fixes its own work.

On Terminal-Bench 2.1, the workflow moved DeepSeek V4 Flash from 67.42% to 82.02%, showing how much performance can come from the workflow around the model. The author cautions gains aren't guaranteed on every task, and the extra loop costs more time, tokens, and money — best suited to difficult or long-running tasks.

Repo: github.com/Spielewoy/autoprompt-skill

Original post →

More from coding & agent

coding & agent channel →