Why LLMs Remember but Fail to Apply After Fine-Tuning
HKUST · hf · 2026-07-13
This paper investigates a common issue in LLM fine-tuning: models rapidly memorize new knowledge but fail to apply it during downstream reasoning tasks.
The authors formalize this phenomenon as the Knowing–Using Gap, characterized by two aspects:
- The accuracy discrepancy between memorization and generalization.
- The time lag between a fact being memorized and becoming practically usable.
To explain the underlying mechanics, the paper introduces a self-patching intervention method. This tracks the spatial diffusion dynamics of knowledge within the model and identifies locations where migrating representations to the correct layers significantly resolves failed cases. The findings support a "knowledge circuit mismatch" hypothesis: although the knowledge exists internally, it is not routed to the layers engaged in computation. Finally, the authors propose a simple heuristic strategy that recovers about 58%–75% of the oracle margin in generalization failures.
Related event: Study Reveals Knowing-Using Gap in LLM Fine-Tuning(2 posts)→
More from Research
- Navier-Stokes, Riemann, P vs NP: what this week's math buzzwords mean for you — koltregaskes · 2026-09-11
- Harry Collins: LLMs can't do frontier science because they can't invent new language — whoamisri · 2026-09-11
- The Waymo effect: how AI is quietly making research less collaborative — JohnHammersley · 2026-09-11
- Causal-only attention for non-generative tasks is wasteful, argues HF engineer — antoine_chaffin · 2026-09-11
- Catholic University of Chile researcher: scaling AI feedback is key to sustainable medical education — julianvarascom · 2026-09-11
- Nature paper images cellular activity across all organs, revealing body-wide circuits — arjunrajlab · 2026-09-11