ASCENT: online test-time training lets LLM agents self-distill verified experience into weights
Haodong Lu · hf · 2026-10-06
- Introduces Online Agentic Test-Time Training (OaTTT): an LLM agent trains its own weights on deployment-time execution trajectories, with the trajectory plus its verification result as the only learning signal, persisting across a stream of tasks.
- Directly imitating or reinforcing single-attempt tokens destabilizes the policy; ASCENT instead self-distills verified experience — a frozen copy of the LLM receives the trajectory as privileged information, predicts next-token distributions with hindsight, and distills them into persistent LoRA fast weights, with no external reference solution or stronger teacher.
- Invalid-action turns are further removed to distill enhanced privileged experience.
- Across ALFWorld, WebShop, and AppWorld at varied model scales, ASCENT improves task success and interaction efficiency as experience accumulates, outperforms online adaptation methods, and transfers to held-out scenes — consolidating verified experience into weights without a separate training phase or memory retrieval.
More from coding & agent
- OpenCode hits 17M MAUs just 16 months after launch — ycombinator · 2026-10-06
- Claude Code 2.1.291 Fixes Session-Drop Regressions, System Prompt Tokens Jump 15% — ClaudeCodeLog · 2026-10-06
- Claude Code 2.1.291 fixes permission-prompt and message-loss regressions — ClaudeCodeLog · 2026-10-06
- Claude Code 2.1.291 fixes cloud session and message-loss regressions — ClaudeCodeLog · 2026-10-06
- AI agent architecture explained: the 7 modules from perception to observability — ZabihullahAtal · 2026-10-06
- Memoria 1.0: a local, model-agnostic LLM memory layer hitting 89.8% Recall@1 on LongMemEval — kitkatz69 · 2026-10-06