A year-built personal agent was finally beaten by a one-day-old competitor
Antony_Richards · reddit · 2026-07-23
- After nearly a year building a personal assistant agent, the author still couldn’t tell whether it was actually good using informal observation alone.
- Existing benchmarks and desktop-model-based evaluation were not helpful because the agent tends to remember wins and forget misses, and performance drifted after model swaps, memory failures, and tool changes.
- The only thing that worked was writing fresh tests the agent had not seen before — more like a heartbeat than a final score.
- A year-old agent was eventually beaten on an entire category by a one-day-old agent, which made the improvement gaps obvious.
- The post asks how other people judge whether an agent is genuinely improving: self-tests, external tests, or just waiting for failures.
More from coding & agent
- PyTorch backs AMD’s AI DevMaster Hackathon with $30,000 in prizes — PyTorch · 2026-07-23
- Warp’s factory foreman agent introduces itself before starting work — vikvang1 · 2026-07-23
- Forrester says agentic AI adopters now prioritize integration over model quality — rseroter · 2026-07-23
- Local open-source agent says it beats Hermes 37 to 31 on GAIA Level 1 — kimmonismus · 2026-07-23
- Local AI agent in DWN.BRIDGE could read files outside its workspace — dwn270787 · 2026-07-23
- Claude Code 2.1.218 adds CLI, MCP, and accessibility updates — ClaudeCodeLog · 2026-07-23