One Layer Is Enough: Single‑Layer RL Beats Full‑Parameter Training Across Models, Tasks, and Algorithms
A systematic study of reinforcement‑learning fine‑tuning for large language models reveals that RL gains are highly concentrated in a few middle Transformer layers, and training just one such layer can match or even exceed full‑parameter RL performance across multiple models, tasks, and algorithms.
