Keep your failed agent tasks — rerun them on each model release to catch the breakthrough

victor_explore · x · 2026-09-30

A practical methodology from @victorexplore: archive every failed agent run and treat them as an eval pile for future models.

Key points:

The post cites an Anthropic Labs anecdote: Mike Krieger said a shelved computer-use agent project was rerun on every new model, and with Claude 3.7 it finally succeeded more often than not. Failed tasks aren't garbage — they're a backlog waiting for the next model to unlock.

Original post →

More from coding & agent

coding & agent channel →