Let Agents Rewrite Their Own Harness — But Keep the Grader Independent
DESI_MOGGER · reddit · 2026-09-12
A Reddit proposal for testing genuine agent self-improvement: let the agent rewrite parts of its own harness (prompts, tool handling, retries) while keeping the grader completely outside the loop. Running against the same independent evaluation after each change makes improvement measurable and prevents the agent from quietly optimizing its own scoring logic instead of actually getting better.
More from coding & agent
- llama.cpp launches llama.app: frontier AI fully local with one command — ngxson · 2026-09-13
- How do production agent systems handle side effects across mixed frameworks? — merlinofthewater · 2026-09-13
- Feeding GT and F1 footage to GPT-6 Astra pushes AI 3D graphics near game-grade — cedric_chee · 2026-09-13
- SWE-2 Draws Praise: Devin Called the Most AGI-Flavored Engineering Harness — silasalberti · 2026-09-13
- threejs-game-skills hits 2k stars: agent skills for polished Three.js browser games — majidmanzarpour · 2026-09-13
- A repair-loop stop rule needs three clauses: budget, evidence, escalation — alexcovo_eth · 2026-09-13