PB-Scale Embodied Data: Efficient Storage & Training with Volcano Engine Multimodal Data Lake

Digua Robot leverages Volcano Engine's multimodal data lake — combining TOS object storage, Lance columnar format, and Ray distributed processing — to unify heterogeneous robot episode data, enable PB-scale ingestion, spatial quality checks via LAS operators, and achieve near-PFS training throughput with RobotLanceDataset.

ByteDance Data Platform
ByteDance Data Platform
ByteDance Data Platform
PB-Scale Embodied Data: Efficient Storage & Training with Volcano Engine Multimodal Data Lake

What a Single Embodied Episode Contains

A robot episode records one complete task execution and includes joint states, actions, multi-view video, camera calibration, timestamps, and task text. Data is organized via two-level indexing: episode for the full task and frame (timestep) for each time point. A single trajectory mixes large video objects with high-frequency time-series state/action data, often originating from different file formats — HDF5, JSON, PKL, raw video — creating infrastructure complexity.

Four Challenges at Million-Episode Scale

1. Heterogeneous Data Formats

No unified standard exists across datasets. Directory structures, field naming, coordinate systems, and control semantics differ. Visual data may be wrapped in HDF5 or stored as separate video+timestamp files; gripper states use 0–1 or absolute joint positions; action can represent target positions or deltas. Adapting training code per format would require continuous maintenance.

2. Storage Cost Pressure

Traditional clusters rely on local disks or parallel file systems (PFS) for read performance, but costs escalate at hundreds of TB to PB scale. Object storage (TOS) offers elastic capacity, decouples data from compute, and enables multi-cluster sharing, making it a better data-lake foundation. However, object storage cannot simply replace local FS; reducing small-file accesses and using batched, concurrent reads are essential for training throughput.

3. Vector Retrieval & Governance Needs

As data grows, semantic search over image, language, and trajectory embedding vectors becomes critical for finding similar episodes, clustering tasks/scenes, detecting duplicates/anomalies, and curating training sets. Scattering metadata, time-series data, and vectors across systems complicates versioning and primary-key alignment. Lance supports Arrow-style columnar data plus vectors and vector indexes, allowing structured data and vector capabilities to be linked via a unified episode/frame ID. The minimal DROID schema used here does not yet generate embeddings directly; future extensions can add separate embedding tables recording episode_index, frame_index, model_version, and embedding without rewriting raw state/video data.

4. Distributed Processing for Million-Episode Ingestion

Ingestion involves object enumeration, JSON/HDF5 parsing, video validation, calibration conversion, and columnar writing. Single-machine serial processing would take days to weeks; failures incur high re-run costs. Episodes are independent, making them ideal for parallelization. Ray distributes conversion tasks across a cluster, handling resource scheduling and task-level retries. Scaling compute resources reduces large-scale conversion from days/weeks to hours.

Digua Robot × Volcano Engine: Unified Embodied Data Foundation

The solution centers on TOS + Lance + Ray : TOS stores raw data and video objects cost-effectively; Lance organizes robot states, actions, camera calibrations, and provides vector-search extensibility; Ray handles distributed conversion of million-episode datasets. Training reads via RobotLanceDataset directly fetch lake data, performing time-window sampling, multi-camera video decoding, and state-action alignment.

Key Design: Separate Video from High-Frequency Time-Series

Videos are transcoded to MP4 and stored as objects in TOS; Lance holds states, actions, calibrations, frame indices, and video URIs. Training decodes video on demand rather than writing decoded frames into the lake. This yields four benefits:

Lower storage cost: MP4 objects avoid massive small-image files and duplication.

Higher training read efficiency: When only states/actions are needed, full video is not read; loading is task-selective.

Clearer governance: Raw data remains in TOS; lake ingestion produces standardized objects without overwriting sources, enabling traceability, updates, and reprocessing.

Flexible evolution: Video transcoding, structured schema, and embedding can evolve independently, leaving room for semantic search, data filtering, and model iteration.

Why Lance as the Unified Embodied Data Format?

Three factors drove the choice: training read efficiency, object-storage native read/write, and vector-search extensibility. Raw robot data (HDF5, JSON, PKL) suits acquisition/archival but not large-scale training — direct reading would require continuous adaptation to directory structures, field semantics, and codecs. Parquet offers mature columnar storage but needs extra handling for cross- row group random access and indexing. Lance natively supports vectors and vector indexes, making it a natural fit for unifying robot states, actions, video references, and multimodal embedding in one system. Crucially, Lance reads/writes directly on TOS : structured data and MP4 videos both reside in object storage, so training clusters need not copy full datasets to local or PFS, further decoupling compute and storage.

The team evaluated LeRobotDataset but, at project start, its standard access paths targeted Hugging Face Hub or local FS, not a TOS-based massive data lake. They therefore wrapped a lightweight RobotLanceDataset around Lance, preserving core capabilities: episode/frame two-level indexing, timestep -based sampling of states/actions/video, direct TOS reads of Lance data and MP4, and multi-camera parallel decoding. As the LeRobot community adopts Lance backends, the technical direction converges; for Digua Robot, Lance is the connective layer linking object storage, data governance, training reads, and vector search.

Two Tables Manage One Robot Trajectory

Using the DROID dataset as an example, lake data splits into frames.lance and episodes.lance. frames.lance stores per-timestep high-frequency time-series: robot states, actions, camera parameters. episodes.lance stores trajectory-level metadata: task description, video URIs. The two tables join on episode_index. This design separates high-frequency time-series, trajectory-level info, and video objects while linking them via a unified index. Training reads only required fields and time windows, avoiding full trajectory or video loads. Adding new state fields, video modalities, or embedding does not alter the existing organization.

Ray-Accelerated Million-Episode Lake Ingestion

After unifying formats and storage, the next challenge: efficiently ingesting million-episode datasets. Ingestion is not simple file copying; each trajectory undergoes parsing, video processing, state-action alignment, calibration conversion, etc. At million-episode scale, serial processing takes days/weeks. Using Volcano Engine's multimodal data lake distributed compute, the team parallelized with Ray. Each episode becomes an independent task dispatched to the cluster, scaling horizontally with added resources.

The ingestion pipeline: Raw Data → Ray Distributed Processing → Format & Index Unification → Write to Multimodal Data Lake → Data Quality Checks .

From Single-Node Conversion to Distributed Ingestion

Ray handles raw-data parsing, video path reading, state-action alignment, camera parameter conversion, finally writing standardized episode/frame data to Lance; videos are transcoded to MP4 and stored as TOS objects. Compared to serial processing, this fully utilizes cluster resources and task-level retries reduce large-scale re-run costs, making continuous million-episode ingestion feasible.

Beyond "Data In Lake" to "Data Usable"

Storing data is insufficient. Video corruption, action-frame misalignment, calibration anomalies can degrade training. Post-ingestion quality checks run distributed via Ray, focusing on video integrity and temporal sync between robot states, actions, and video. For deeper spatial quality issues, the team employs Volcano Engine's LAS Spatial Perception Operators to validate camera intrinsics/extrinsics. Example: forward kinematics (FK) on the robot model combined with camera intrinsics/extrinsics projects the manipulator into the image; comparing visual segmentation with the projection reveals extrinsic offsets, coordinate-definition errors, or camera misconfigurations. In short, traditional checks verify "file intact, time aligned"; LAS verifies "robot 3D spatial data is correct". Results are written back as quality fields in Lance for downstream anomaly filtering, root-cause tracing, and training-set construction. Thus the multimodal data lake covers distributed ingestion, data quality detection, and spatial data governance, ensuring stable, usable data for training.

Training Performance: Can Object Storage Match PFS?

The team designed RobotLanceDataset for efficient reads: structured data fetched on-demand via Lance, video retrieved by index from TOS and decoded, aligned on the training side. For the common "multi-camera at one timestep" scenario, each video opens once and batches decode required frames, reducing repeated object-storage access and decoder re-initialization overhead.

End-to-end chain: TOS/Lance Data Lake → On-Demand State/Action Reads → Batched Multi-Video Reads → Time-Window Alignment → Model Training .

Dataset Initialization: 27.5 s → 0.02 s

Benchmarking RobotLanceDataset against LeRobotDataset local reads on 100 episodes with three camera streams: initialization dropped from 27.536 seconds to 0.020 seconds . Lance's organization means startup only opens the dataset and reads necessary episode metadata — no full scan or large cache warm-up. Per-sample read latency remains comparable, showing no training-phase throughput sacrifice.

TOS vs. PFS: Throughput Comparison Under Realistic Loads

Simulating real training workloads, the team compared TOS and PFS. As per-read data volume grows, TOS throughput rapidly approaches PFS. Sequential read of 250 timesteps: TOS achieves 95.3% of PFS throughput. Under random sampling (closer to actual training), TOS sustains 90%+ of PFS across read sizes. In typical embodied-training batch and random-access patterns, TOS delivers near-PFS performance while retaining object-storage elasticity and cost advantages. The key is not raw media speed but the TOS + Lance data organization plus batched sampling and multi-video parallel decoding minimizing remote-access and decode overhead.

Sequential 250 timesteps: TOS reaches 95.3% of PFS Random sampling: TOS reaches 90%+ of PFS

From "Store Data" to "Use Data": A Full-Stack Embodied Multimodal Data Lake

As robot data scales from TB to PB, infrastructure must go beyond "store it somewhere". From heterogeneous unification and low-cost storage to million-episode processing, quality governance, and model training, each stage demands new capabilities. Digua Robot's team built a full-lifecycle embodied data foundation on Volcano Engine's multimodal data lake:

Unified Storage & Organization: TOS + Lance ingest video, robot states, actions, camera params into one lake.

Scalable Data Processing: Ray parallelizes million-episode conversion/ingestion for high throughput.

Embodied Data Quality Governance: LAS spatial operators detect calibration, spatial-relation issues unique to embodied data.

Training-Efficient Reads: RobotLanceDataset, batched sampling, multi-video parallel decoding let lake data directly serve model training.

Tests confirm: with object-storage cost/elasticity benefits, TOS hits 95.3% PFS throughput in large sequential reads and 90%+ in random sampling. For Digua Robot, this solves not a single storage/compute point but creates a complete ingestion–processing–governance–training pipeline. As embodied AI advances toward larger scales and richer modalities, data infrastructure must evolve from "fits" to "well-managed, fast-processed, readily usable" — the core problem this practice addresses.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

embodied AILance formatmultimodal data lakedata quality governancePB-scale dataRay distributed processingrobot training dataTOS object storage
ByteDance Data Platform
Written by

ByteDance Data Platform

The ByteDance Data Platform team empowers all ByteDance business lines by lowering data‑application barriers, aiming to build data‑driven intelligent enterprises, enable digital transformation across industries, and create greater social value. Internally it supports most ByteDance units; externally it delivers data‑intelligence products under the Volcano Engine brand to enterprise customers.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.