Scaling LLM RL: How Batch Size Affects Training Speed and Efficiency
The article explores how batch size scaling impacts LLM reinforcement learning training efficiency, introducing critical generation and training batch concepts, showing optimal batch size balances sample efficiency and system throughput, with experiments achieving 29% time reduction on fixed hardware.
Introduction: Efficiency as the Key Variable in Scaling RL
As pretraining brings large models to new capability starting points, reinforcement learning (RL) drives further reasoning and action abilities through continuous exploration and feedback. The larger the role of RL in model development, the more directly training efficiency impacts compute budgets. A illustrative calculation: 10,000 GPUs running for 30 days at $3 per GPU-hour costs $21.6M. A 10% reduction in training time saves 720,000 GPU-hours, or ~$2.16M — directly reducing cloud bills or freeing capacity for the next experiment cycle.
These efficiency gains hinge on a fundamental parameter present in almost every training recipe: Batch Size . In practice, batch size appears in two decisions:
Batch Size Tuning : Adapting the recipe to new conditions (model size, GPU type). Too small a batch underutilizes hardware; mismatched learning rates slow or destabilize learning.
Batch Size Scaling : Converting additional parallelism into shorter wall-clock time. Given more GPUs, can batch size scale proportionally so the model reaches target performance sooner? Even with fixed GPUs, if generation concurrency is low, can a larger batch utilize idle capacity?
Both decisions ultimately ask: How to find the optimal batch size that minimizes wall-clock time to target performance?
Larger Batch: Accelerator or Brake?
Tencent Hunyuan team experiments show that moderately increasing batch size, with coordinated learning rate adjustments, can maintain learning efficiency while improving throughput. However, excessively large batches cause the extra sample cost to outweigh throughput gains. Figure 1 illustrates this as a mountain climb: some routes reach the same performance peak faster; others appear to take larger steps but actually summit slower. The normalized time-to-target metric captures this trade-off.
Figure 1: Training as mountain climbing. Peak = target performance. Four routes = different batch sizes. Markers = optimizer updates. With learning rate scaling, B, 2B, 4B maintain similar sample-level progress; larger batches reach peak faster due to higher throughput. At 16B, sample efficiency loss exceeds throughput gain, making it slower than B.
From Classic Batch Scaling to LLM RL
The question of how far batch size can effectively scale is not new. In 2017, Goyal et al. (Facebook) demonstrated linear learning rate scaling and gradual warmup, training ResNet-50 on ImageNet in 1 hour on 256 GPUs (batch 8192) vs 29 hours on 8 GPUs (batch 256), preserving accuracy. This shifted the question from "can large batch train?" to "how far can it scale effectively?"
In 2018, McCandlish et al. (OpenAI) linked diminishing returns to gradient noise scale. Shallue & Dahl (Google Brain) noted data parallelism benefits are task-dependent and fair comparison requires per-batch hyperparameter tuning. Related is Batch Size Invariance : after adjusting learning rate, does the model achieve similar learning trajectories when consuming the same number of samples? Hilton et al. (OpenAI, 2022) studied this for policy optimization.
What Makes LLM RL Different?
In supervised learning, training samples come from a fixed dataset. Larger batch mainly adds forward/backward/update cost. Classic batch scaling discusses how larger batch changes gradient estimation, update count, and parallelism.
LLM RL adds a critical loop: the Transformer acts as an agent that must first autoregressively generate responses/trajectories or interact with environments to obtain rewards; these online-generated data then enter the learning phase for forward/backward/update. The updated Transformer then generates the next batch. The model is both the inference system producing training data and the learning system consuming it.
Architecturally, generation (autoregressive decoding, repeated weight reads, KV cache, concurrent sequence scheduling) and training (batched forward/backward, activation storage, gradient computation, optimizer updates) use GPUs differently. Scaling batch size is not merely feeding more data to the same computation: it can improve generation efficiency via higher concurrency while increasing training compute and memory overhead. Learning curves alone cannot answer how much faster training will be; generation throughput alone cannot answer how many more samples the model needs.
Two Efficiencies, One Analytical Framework
The ultimate goal is not maximum batch size or fastest local stage, but earliest achievement of required capability. The paper uses Time-to-target : real time from start until model first hits a preset validation target.
Time-to-target is jointly determined by:
Sample efficiency : learning progress per training response. Batch size and its hyperparameters change learning effect under fixed sample budget.
System efficiency : end-to-end pipeline throughput (responses generated and processed per second), including generation, update, waiting, and synchronization.
Let J be validation performance, N cumulative training responses, t real time. Along a smooth learning trajectory:
If throughput doubles but per-response progress halves, they cancel. For actionable comparison in noisy, discretely evaluated experiments, the paper fixes a validation target. Let N* be responses needed to reach target, q be average end-to-end throughput during that period. Then time-to-target T = N* / q.
Relative to a reference config, let ρ = N*_new / N*_ref (sample cost, >1 means more responses needed) and σ = q_new / q_ref (throughput gain, >1 means more responses/sec). Then T_new / T_ref = ρ / σ.
This yields two perspectives:
Theoretically : compare throughput gain σ vs sample cost ρ. Only when σ > ρ does larger batch reduce training time.
Practically : first establish invariance, then optimize throughput. Re-tune batch-dependent hyperparameters to find a batch range where learning curves (plotted against cumulative samples) approximately align. In this range, ρ ≈ 1, so improving throughput directly reduces time-to-target. If system tuning changes learning behavior, re-verify alignment.
Step 1: Preserve Per-Sample Learning Effect
Batch Size Invariance means: if batch doubles, optimizer updates to reach same performance should roughly halve, so total sample demand stays constant. After per-batch hyperparameter tuning, a fair comparison should show: more samples yield better performance; at equal cumulative samples, different batches achieve similar performance.
This does not happen automatically: larger batch averages gradients over more responses but reduces update opportunities under fixed sample budget. To maintain per-sample learning effect, each update must step further. For Adam, the team uses the square-root rule as learning rate starting point:
This rule is only a starting point; only learning rate is changed, other optimizer hyperparameters fixed. Joint tuning of more hyperparameters may improve invariance further (left for future work).
For clean comparison, experiments use a custom training system evaluating GRPO and PPO. Each rollout batch receives a single global optimizer update — no minibatching, no rollout reuse — avoiding extra tuning dimensions like minibatch size and optimization epochs.
Finding 1: GRPO Approximate Invariance over 16× Prompt Batch Range
GRPO training on Qwen3-30B-A3B-Instruct-2507, filtered DAPO-MATH-17K data, evaluated on AIME 2024/2025 mean@32. A training sample = filtered response retained for optimizer update. One update contains P prompt groups, each with G responses, so batch size B = P×G. With G=8, learning-rate-adjusted runs from P=64 to P=1024 show similar sample-level learning curves; P=2048 and P=4096 fall below this family.
Right panel shows actor gradient norm roughly scales as 1/√B in the invariant range; at P=2048/4096 the decline flattens, coinciding with learning curve misalignment. This matches diminishing marginal returns of gradient averaging, though gradient norm is only an indirect diagnostic.
Finding 2: PPO Actor and Critic Have Different Scales
PPO uses internal Hunyuan policy and value models (3B active params each), mixed math/science/logic data, average benchmark score. G=1, so B=P. Batches 256–2048 align; 4096 falls below.
Actor gradient norm declines continuously, flattening at largest batch; critic gradient norm stays relatively flat with larger variance. Different training objectives may imply different effective batch scales, suggesting future separate analysis and tuning.
Finding 3: For GRPO, Total Response Batch Matters More Than Prompt Count or Group Size Alone
GRPO batch can scale via more prompts (P) or more responses per prompt (G). Larger G reduces gradient variance via averaging, but responses within a group share prompt and construct advantage via group-relative rewards — not fully equivalent.
Configurations with same total responses (P,G) = (128,8) vs (64,16) [1024 responses], and (256,8) vs (128,16) [2048 responses] show similar learning trajectories (left) and similar actor gradient norms (right). Gradient norm for 2048 responses is consistently lower than for 1024. This supports using B = P×G as GRPO batch unit in tested configs, but does not imply P and G are arbitrarily interchangeable.
Finding 4: Keeping Learning Rate Fixed Breaks Approximate Invariance
In practice, adding GPUs often expands data parallelism and global batch. If learning rate stays at small-batch setting, training may not diverge and can appear more stable due to lower gradient noise. But with fewer updates for same responses, model may not reach same performance. Training stability ≠ Batch Size Invariance.
Reference: P=128. Compare two P=256 runs: one with lr scaled by √2, one with fixed lr. Scaled-lr run tracks reference; fixed-lr run lags significantly (e.g., at ~120k responses: 74% vs 77%). Ideal invariance predicts update count ratio 0.5×. Scaled-lr achieves 0.50–0.67×; fixed-lr requires 0.75–0.83× updates, which translates to 1.50–1.67× sample consumption due to double batch size.
Step 2: Improve Throughput on Fixed Hardware
After per-sample learning effect is approximately preserved, next question: how to process samples faster? Common approach: add GPUs, expand data parallelism. But on fixed hardware, can larger batch accelerate? Opportunity comes from compute asymmetry between generation and training .
Generation : Autoregressive decoding produces one token per sequence per step. Low concurrency means repeated weight reads for few token vectors. Increasing active sequences amortizes weight-read overhead, so response count may grow faster than generation time.
Training : After responses generated, forward/backward processes many token positions in parallel, already exposing high parallelism. Larger batch mainly adds work. In measured GRPO configs, doubling prompt batch from 128 to 256 increased actor update time from 101.3s to 208.6s — roughly linear with batch.
A simplified local model captures this asymmetry:
Where B_gen = responses in logical rollout batch, B_train = retained responses per optimizer update. Experiments use asynchronous partial rollout with B_gen ≥ B_train, so scaling training batch also scales generation batch. α represents generation overhead amortizable over concurrent responses; training time scales near-linearly with batch.
Roofline perspective: batching changes arithmetic intensity and whether execution is memory-bandwidth or compute bound. The simplified model summarizes this generation-vs-training cost difference.
Finding 5: On Fixed Hardware, Generation Time Grows Sublinearly
Figure 7 shows response count and collection time relative to smallest batch. Blue (responses) grows faster than orange (time); ratio = generation throughput gain. PPO: batch 256→1024 (4× responses), collection time 39s→68s → throughput gain 4/(68/39)=2.29×. GRPO: P=128→256 gain 1.31×; 128→512 gain 1.36×. Gain from 256→512 is small, indicating saturation of exploitable generation throughput in tested range.
Stage conclusions:
LLM RL generation and training have compute asymmetry: generation limited by weight reads/memory bandwidth; training already high parallelism, time ~linear with batch.
Larger batch lets more concurrent decoding sequences amortize weight-read cost, potentially improving generation throughput even on fixed hardware.
Step 3: Combine Two Critical Batches to Find Minimum Time-to-Target
Results reveal two distinct boundaries in LLM RL batch scaling:
Critical generation batch (system efficiency): For fixed model, decoder, response length distribution, hardware allocation, marks where generation throughput plateaus. Before it, larger batch improves hardware utilization; after, little extra throughput.
Critical training batch (sample efficiency): Under current hyperparameter tuning, upper bound of approximate Batch Size Invariance. Before it, larger batch needs similar responses to target; beyond, marginal response benefit no longer compensates fewer updates, so required responses increase.
They are separate because they answer different questions: how much generation throughput remains, and whether training can maintain per-response learning. Memory capacity may limit feasible batch before throughput plateaus. Neither boundary alone gives fastest config: if generation and training batch scale together along same B, their relative order shapes time-to-target curve via ρ/σ balance.
Three regimes: 1. Learning preserved, throughput rising → time decreases. 2. Learning preserved, throughput saturated → flat time region. 3. Extra sample cost dominates → time increases. Order can reverse: if learning degrades first but throughput still rises fast, larger batch may still be faster as long as σ > ρ. Goal is minimize time-to-target, not mechanically pick one critical batch.
Measured Results: Best Config Cuts Time 29%, Worst Fixed-LR Slows 42%
With learning rate rescaled, increasing P from 128 to 512 and 1024 reduces normalized time-to-target to 0.74× and 0.71× . At P=1024, same 122.88K retained responses reach target, end-to-end throughput +41%, training time drops from 11.90h to 8.42h — 29% faster without adding GPUs .
But larger batch not always better. At P=2048 and P=4096, required responses increase ~60%, exceeding 30% and 36% throughput gains, time rises to 1.23× and 1.18× . Even at moderate P=256, fixed learning rate yields only 17% throughput gain but 67% more responses needed, time becomes 1.42× reference. Outcome always hinges on whether throughput gain compensates extra sample cost.
Practical Two-Step Tuning Framework
First, align learning effect : For each candidate training batch, adjust learning rate, compare learning curves at equal cumulative responses, estimate sample cost ρ.
Then, compare system gains : Within memory limits, adjust generation batch or decoding concurrency, measure end-to-end throughput q, finally select config minimizing ρ/σ, confirming it reaches target earlier.
Significance and Boundaries
This work does not propose a new RL algorithm but provides a simple, scalable, hardware-aware thinking framework for existing GRPO/PPO training. Four takeaways:
Theoretical understanding : Separates two often-conflated questions — can model learn equally well with same samples? Can samples be processed faster on given hardware? LR scaling maintains approximate invariance within a range, but the law eventually breaks. Predicting this effective range could reduce per-model/task re-tuning cost.
Training practice : Verify learning first, then speed. When changing GRPO/PPO training batch, first adjust LR, check if equal cumulative responses yield similar performance; then increase generation concurrency, measure true end-to-end throughput. Reuse tuning sequence across model/task/hardware changes, but do not directly copy previous batch numbers.
System design : Distinguish generation and training costs before deciding how to scale batch. Generation limited by weight reads/bandwidth — higher concurrency amortizes this; training has higher arithmetic intensity — larger optimizer batch may not yield proportional gains. Decoupling generation and training batch, allocating GPUs accordingly, may further unlock throughput, but must avoid excessive staleness breaking learning behavior.
Evaluation methodology : Different batches must start on same starting line. Each group tunes its own batch-dependent config first, then answer: with same samples, who learns better? To same capability target, who takes less time? Comparing only same update steps may mistake "saw more samples" for "learned better"; comparing only throughput may mistake "processed faster" for "finished training earlier".
Boundaries are clear: when model, reward, task distribution, training stage, inference engine, or execution strategy change, both sample efficiency and system throughput may shift; previously measured effective batch ranges cannot be directly reused. Natural next steps: go beyond "only LR tuning" to find transfer rules for other batch-dependent hyperparameters; PPO experiments hint actor and critic may need different batch scales.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
