New Study: Training a Single Transformer Layer Can Match Full-Parameter RL

tokenbender · x · 2026-08-16

arXiv paper 'Is One Layer Enough?' finds that training a single transformer layer can recover most of the gains of full-parameter RL training, sometimes surpassing it. Across 7 models, 3 RL algorithms, and multiple task domains, high-contribution layers concentrate in the middle 40-60%.

Original post →

More from Research

Research channel →