A unified learning-dynamics view ties data attribution, forgetting, and plasticity loss to token interactions
Yi Ren · hf · 2026-10-01
This paper derives a token- and layer-wise decomposition of how learning one token changes another prediction, separating softmax force, shared readout geometry, and residual connections. Positive interactions identify useful experience; negative ones cause concentrated collisions or accumulated erosion; over time updates reshape the readout geometry that mediates future learning. This single evolving interaction unifies data attribution, forgetting, and plasticity loss, yielding practical data selection, interference controls, and a readout-based diagnostic of future learnability.
More from Research
- Mathematician details agentic math workflow with Codex CLI and Claude Code, results coming — burny_tech · 2026-10-01
- Egocentric video dataset EgoPro from LightwheelAI trends on Hugging Face — LightwheelAI · 2026-10-01
- NVIDIA Opens 2027–28 Graduate Fellowships With $60,000 Awards, Due Oct 30 — NVIDIAAI · 2026-10-01
- Airbench Crowdsources a Local LLM Leaderboard via One-Prompt Agent Benchmarks — dh7net · 2026-10-01
- Five UW Allen School Faculty Win NSF CAREER Awards Spanning Quantum Crypto to Brain-Inspired AI — lazowska · 2026-10-01
- GDP.xlsx benchmark: 70 real spreadsheet tasks, best frontier agent scores only 38.3% — echen · 2026-10-01