Keep your failed agent tasks — rerun them on each model release to catch the breakthrough
victor_explore · x · 2026-09-30
A practical methodology from @victorexplore: archive every failed agent run and treat them as an eval pile for future models.
Key points:
- Every broken agent run is effectively an eval sample.
- Rerun the whole pile on each new model release to catch the exact week the model suddenly handles your workload.
The post cites an Anthropic Labs anecdote: Mike Krieger said a shelved computer-use agent project was rerun on every new model, and with Claude 3.7 it finally succeeded more often than not. Failed tasks aren't garbage — they're a backlog waiting for the next model to unlock.
More from coding & agent
- Building a Code Review Agent That Learns From Feedback With Groq and Hindsight — pasulabhavya · 2026-09-30
- Open-Dots, an open-source clone of OpenAI's Dots, hits 4,500 GitHub stars in 24 hours — matchaman11 · 2026-09-30
- Graphsub pitches in-memory graph DB for agent data: don't trust one AI corp with it all — arthurcolle · 2026-09-30
- Ex-Googler: AI agents can't write prod code yet, but what else fits a 60-minute interview? — prajdabre · 2026-09-30
- Open-Source Offline AI Speaking Coach Built With Ollama and faster-whisper — iamrishavraj1 · 2026-09-30
- Recoverable shared memory for multi-agent setups: files, append-only logs, provenance — RocketSeven · 2026-09-30