Self-generated feedback destabilizes test-time training, causally decomposed at 128K tokens
KAUST · hf · 2026-10-06
Test-time training (TTT) lets models write into their weights during inference, but learning from the model's own output creates a feedback loop: each update changes the model generating the next training example. The paper causally decomposes this failure:
- Across 128K-token streams, retaining generated-text updates worsens prediction on independent human text in three TTT-E2E configs (125M/760M/3B) and when Adam updates Qwen3-4B weights
- Fixed Generation (frozen model writes training chunks) removes >98% of the damage; Recorded Replay separates loss from reading degraded text vs. updating on it; paired one-update comparisons show updates predict their source better but fresh real text worse
- The cost compounds under Closed Loop adaptation, with few trajectories causing most large failures
- Settlement checks candidate states on independent real text before committing, keeping endpoint gaps at 0.07/-0.02 nats while retaining real-text adaptation
Writing itself isn't the failure—updating on your own updates is. Check predictions on independent evidence before retaining updates.
More from Research
- O(n²) matrix multiplication is almost certainly false even if ω = 2, says basedjensen — basedjensen · 2026-10-06
- Cohere Labs Releases Tiny Aya, a Family of Small Models Covering 70+ Languages — Cohere_Labs · 2026-10-06
- Will LLMs wreck elegant math? O(n²) limits are already crumbling — teortaxesTex · 2026-10-06
- ETH's AME-2 legged locomotion paper accepted at TRO, training code open-sourced — ChongZzZhang · 2026-10-06
- 0.8B model beats 2B on ARC-Challenge (42.15%) via closed-form weight surgery with zero backprop — AdventurousTwo6445 · 2026-10-06
- MIT team's Science Task Taxonomy maps 232 subfields and 208,202 scientific tasks — JMateosGarcia · 2026-10-06