Single-Layer Training Matches Full-Parameter RL

jiqizhixin · x · 2026-07-15

A preprint from researchers at the University of Minnesota, Peking University, and Amazon suggests that training just a single Transformer layer can match or even outperform full-parameter reinforcement learning fine-tuning.

The method identifies the few critical intermediate layers responsible for most RL gains and updates only those. Experimental results show it outperforms full-parameter RL on math, code, and agentic tasks. The paper concludes that layer specialization might be the key to performance improvements.

Original post →

More from Research

Research channel →