Why LLMs Remember but Fail to Apply After Fine-Tuning
HKUST · hf · 2026-07-13
This paper investigates a common issue in LLM fine-tuning: models rapidly memorize new knowledge but fail to apply it during downstream reasoning tasks. The authors formalize this phenomenon as the **Knowing–Using Gap**, characterized by two aspects: - The accuracy discrepancy between memorization and generalization. - The time lag between a fact being memorized and becoming practically usable. To explain the underlying mechanics, the paper introduces a **self-patching** intervention method. This tracks the spatial diffusion dynamics of knowledge within the model and identifies locations where migrating representations to the correct layers significantly resolves failed cases. The findings support a "knowledge circuit mismatch" hypothesis: although the knowledge exists internally, it is not routed to the layers engaged in computation. Finally, the authors propose a simple heuristic strategy that recovers about 58%–75% of the oracle margin in generalization failures.
Related event: Study Reveals Knowing-Using Gap in LLM Fine-Tuning(2 posts)→
More from Research
- LTX 2.3 LoRA demo changes a video’s camera angle — CQDSN · 2026-07-21
- OpenForecaster uses daily news to improve language-model forecasting — Cohere_Labs · 2026-07-21
- SenseTime unveils U1 Pro and open-sources a 50M-sample vision dataset at WAIC 2026 — 机器之心 · 2026-07-21
- Baseten study finds new facts in LLM weights are fragile unless trained from many restatements — alex_verem · 2026-07-21
- Kimi K3 and Fable 5 now look much closer than the old open-vs-closed gap — FinanceYF5 · 2026-07-21
- uv-scripts/ocr returns to the top of Hugging Face datasets with a JSON model picker — vanstriendaniel · 2026-07-21