GPU Starved? DataLoader Optimizations Cut Embodied Model Training Time 2x
Baidu Baige optimizes DataLoader across sampling, indexing, decoding, and post-processing for embodied model training, reducing FastWAM single-sample video processing from 133ms to 63ms, stabilizing indexing at 0.61ms across 27,500 files, and boosting Qwen3-VL-30B throughput 1.68x via load-balanced sampling.
Training embodied models like Vision-Language-Action (VLA) and World Action Models (WAM) relies heavily on video data that cannot be fully preprocessed and stored. Instead, video frames must be read on-demand and decoded online during training. If data preparation cannot keep pace with GPU computation, the GPU idles waiting for data — a waste that accumulates over hundreds of thousands of steps.
1. Batch Formation and the Embodied Model Challenge
Each training step consists of GPU-side compute (forward, backward, update) and CPU/I/O-side data preparation: fetching samples from disk, decoding into tensors, post-processing, and collating into a batch. This pipeline is organized by the DataLoader. The goal of DataLoader optimization is not to minimize batch formation time absolutely, but to ensure it stays below the single-step compute time so that GPU compute fully hides data loading latency.
Large language models tokenize text offline, making data loading lightweight. Vision-Language Models (VLMs) with image inputs add image decoding but remain similar. In contrast, VLA and WAM samples are temporal windows randomly sampled from episodes. A single episode yields hundreds of overlapping windows, which are not pre-sliced and stored because that would explode storage and fix sampling strategies. Instead, raw episodes are kept, and during training a random window is chosen, keyframes located, and frames decoded and processed frame-by-frame online.
2. Optimization Target and Diagnostic Method
Batch size is set by the training side to saturate GPU memory and utilization, fixing the per-step compute budget. Data-side optimization then aims to push actual batch formation time under that budget. The DataLoader pipeline is decomposed into five stages: Sampling (which samples, order, rank assignment), Indexing (mapping sample IDs to file/offset/frame ranges), Decoding (compressed bytes to tensors, the most I/O- and CPU-intensive stage), Post-processing (reshaping, normalization, frame sampling), and Tensor Collation (padding and stacking). The first four stages are analyzed in detail; collation relies on framework defaults.
Optimization proceeds in two cost-ordered steps: (1) tune DataLoader parameters (num_workers, prefetch_factor, pin_memory, persistent_workers) to see if adding CPU cores or storage bandwidth suffices; (2) if bottlenecks remain, profile each stage and apply targeted optimizations.
3. Stage-by-Stage Optimizations on Baidu Baige (hpas.lgn7ib instances)
3.1 Sampling: Load Balancing Across Ranks
In multi-GPU data-parallel training, step time is dictated by the slowest rank. For Qwen3-VL-30B multimodal SFT on LLaMA-Factory (single-node 8-GPU), the default global random sampler assigned widely varying sequence lengths and visual token counts per rank, causing load imbalance. A custom load-balanced sampler was designed: (1) sort samples descending by sequence length to minimize padding within a global step; (2) assign each sample to the rank whose total cost increases the least, while enforcing equal sample counts per rank; (3) shuffle physical rank IDs and global step order to prevent persistent high-cost assignment. Result: training throughput rose from 27.14 to 45.48 samples/s (1.68× speedup).
3.2 Indexing: Avoiding Linear Growth with File Count
VLA/WAM samples are defined by (trajectory ID, start time, window length), requiring a custom mapping to (file, row range, video timestamp interval). FastWAM uses the RoboTwin dataset in LeRobot v2.1 format: one Parquet file per trajectory, one video file per camera. The original implementation used Hugging Face datasets' select interface per sample, which constructs a dataset view each time; view construction overhead scaled with file count. When the dataset grew from 1 to 27,500 files (~6.07M rows), single-index latency jumped from 0.90 ms to 81.1 ms. Switching to the library's built-in fancy indexing (direct positional lookup) bypassed view construction, stabilizing latency at ~0.61 ms regardless of file count. This also illustrates how file organization impacts indexing: LeRobot v3.0 merges multiple trajectories into fewer Parquet files, reducing file-count-induced overhead at the storage layout level.
3.3 Decoding: Reducing Fixed Overhead and Decoding Only Needed Frames
Video decoding cost per call splits into a fixed cost (building decoder object + seeking to keyframe) and a variable per-frame decompression cost. For random access, the fixed cost dominates when multiple cameras or segments require separate decoder instantiations.
Decoder library selection: Benchmarked decord vs torchcodec on UR5e training videos. Per-frame decompression was similar (9 frames: 86.25 ms vs 73.05 ms), but fixed cost per decoder construction differed drastically: decord 141.04 ms vs torchcodec 2.54 ms. Migrating to torchcodec requires pixel-consistency validation (approximate seek may land on different frames for irregular timestamps) and binary compatibility checks with PyTorch versions.
Decode only required frames: FastWAM's visual window contains 33 frames, but only 9 uniformly sampled frames feed the visual encoder. Originally, all 33 frames were decoded, post-processed, then 24 discarded. By passing the downstream 9-frame indices to the decoder, only those frames are decoded. Decoding time dropped from 111.58 ms to 50.97 ms (2.19× speedup). The theoretical linear speedup (33→9 frames) is 3.67×; the gap is due to the fixed decoder construction/seek cost that does not shrink with fewer frames.
Merge repeated decoder calls for the same video: A new model design (inspired by Pi0.7) uses three temporal segments (past, current, future) per camera. The original DataLoader invoked decoding per (segment × camera): 3 segments × 3 cameras = 9 calls per sample, each reopening the same video file and rebuilding the decoder. Merging the frame lists per camera into a single decoder call reduced calls to 3 (one per camera). Without OS page cache: single-sample decoding fell from 77.03 ms to 49.06 ms (1.57×). With cache: ~9% improvement (1.10×), since file open cost was already cached, leaving only decoder reconstruction savings.
3.4 Post-processing: Eliminating Work on Discarded Frames
After the "decode only 9 frames" optimization, the downstream post-processing pipeline still expected 33 frames. To avoid rewriting downstream code, the decoder padded the 9 frames back to 33 (replicating frames), causing per-frame operators (dtype conversion, resize, tensor rearrangement) to run on 33 frames, only to have 24 dropped by the final frame-sampling operator. Since per-frame operators are frame-independent, frame sampling can be moved before them. Combined with the decoding change, the decoder now outputs exactly 9 frames, no padding occurs, and post-processing runs only on those 9 frames. End-to-end video processing latency for FastWAM dropped from 133.0 ms to 63.3 ms (2.10×). Incrementally: decode 9 frames but keep 33-frame padding → 73.5 ms (1.81×); remove padding → 63.3 ms (2.10×).
4. Conclusion: Don't Let GPU Wait for Data
Model and dataset file layouts shift DataLoader bottlenecks. The key is identifying the exact stage that makes the GPU wait. All optimizations above were validated on domestic GPU instances. Baidu Baige continues to leverage systematic AI Infra optimization and highly scalable cluster architecture to deliver cost-effective compute solutions for embodied intelligence scenarios.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Baidu Intelligent Cloud Tech Hub
We share the cloud tech topics you care about. Feel free to leave a message and tell us what you'd like to learn.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
