Research: Diagonal Linear Networks Precisely Characterize Fine-Tuning

ClementineDomi6 · x · 2026-07-08

This study uses diagonal linear networks as a solvable model to precisely characterize fine-tuning behavior, filling the theoretical gap on how pretraining impacts fine-tuning.

Using replica theory, the authors derive exact generalization curves, linking initialization, data structure, and sample efficiency. They show that pretraining hyperparameters shape the model's inductive bias, determining whether fine-tuning reuses pre-trained features or learns new ones. This allows them to distinguish four fine-tuning regimes, including a previously overlooked one that achieves both 'reuse and refinement.'

The analysis extends to imperfect pretraining scenarios: on ResNet and Transformer architectures, the authors prove that when tasks rely on a sparse subset of pre-trained features, a negative relative scale can improve generalization. The core conclusion is that fine-tuning behavior is deeply influenced by the inductive bias inherited from pretraining.

Related event: ICML Paper: How Pretraining Shapes Inductive Bias in Fine-Tuning(5 posts)→

Original post →

More from Research

Research channel →