Why Most Agents Never Learn: The Fix Lives Where the Miss Wasn't
blaizedsouza · x · 2026-10-02
核心观点
多数 agent 在被使用过程中并没有真正变好,因为“修复”所处的位置和当初“失误”的位置不在同一处,自学习闭环必须刻意弥合这个缺口才算有效。
两种学习信号
- agent 自身的 trace(自己的执行记录)
- 用户的真实行为:点击、修正、工具调用、失败
大多数团队只保存 trace,而把人工修正这一半信号完全丢掉,改进信息就此流失。
三层落地位置
捕获到的经验要存放在某处,作者把模型外围分为三层:
- weights:重训模型,成本最高
- in-context:放进上下文,占用窗口
- harness:工程框架层,是团队真正拥有、无需重训就能塑形的一层
一个修正变三条记忆
以一笔被标记的退款为例:它同时变成一条事实(2000 美元上限)、一个案例(这笔被弹回)和一条规则(批准并打标)。单一用户修正在 harness 层可以被结构化为多种知识复用。
More from coding & agent
- Dev says Claude Code with Opus 5.5 is 'a lot of fun': 'It was mostly in my head' — Angaisb_ · 2026-10-02
- NYT editor: AI excels at code for the exact reason it's mediocre at writing — dylfreed · 2026-10-02
- AI shifts game dev from writing code to testing and directing, vets say — AIandDesign · 2026-10-02
- Agent-Reach, Patchright Enhanced, Scrapling: 3 tools to let agents scrape any website — FinanceYF5 · 2026-10-02
- 3 open-source tools give AI agents web-scraping superpowers, even on API-less sites — FinanceYF5 · 2026-10-02
- Claude Opus 5.5 directed a 5.5-minute AI short film from one prompt in ComfyUI — Cheap_Credit_3957 · 2026-10-02