Can 50,000 Web‑Crowdsourced Trajectories Really Strengthen Robot Models? AXIS Benchmark Answers
AXIS demonstrates that web‑based crowdsourced teleoperation data, when systematically generated, cleaned, and augmented, can scale from 50 k to over 1.5 M robot manipulation trajectories, yielding consistent performance gains on the LIBERO‑Plus benchmark and highlighting the importance of task coverage, diversity, and quality control.
Platform and Motivation
AXIS provides a web‑based teleoperation interface that lets participants control a Franka Research 3 arm using keyboard, mouse, or gamepad in a browser. Each interaction generates a human demonstration trajectory without requiring physical robots or local simulation installation.
The approach addresses three challenges: (1) continuous generation of new tasks, objects, and scenes; (2) conversion of noisy community demos into stable, trainable data; (3) verification that dataset growth yields genuine model performance gains rather than merely adding more trajectories.
Data Engine Architecture
AXIS links automatic task generation, browser‑based teleoperation, data cleaning, simulation‑based augmentation, model training, and standardized evaluation into a closed loop.
Initial snapshot (fixed on the Franka Research 3 arm) contains 207 tasks, 50,129 human trajectories, and >60,000 task/scene variants. As of 24 July 2026 the live engine reports 1,816 tasks, 1,530,275 trajectories, and 13,645 h 7 min of data, updated hourly.
Data Cleaning Pipeline
Raw crowdsourced trajectories undergo:
Success‑state verification
Anomaly filtering
Removal of stationary segments
Trajectory smoothing
Fixed‑frequency resampling
Smoothing and resampling reduce average acceleration by 63.9 % and jerk by 80.8 %, substantially lowering high‑frequency noise from web teleoperation.
Simulation Augmentation
Cleaned trajectories are replayed in Isaac Sim where random variations are applied to camera view, lighting, texture, background, object pose, mass, and friction while preserving task semantics and success conditions. This expands the training distribution without altering the underlying task.
Experimental Evaluation
Experiments used the π0.5 base model pre‑trained on different fractions of AXIS data, then fine‑tuned on LIBERO with identical configurations and evaluated on the LIBERO‑Plus robustness benchmark.
Baseline (no AXIS pre‑training) achieved 83.9 % overall success. Pre‑training on 25 % of AXIS data raised success to 84.7 %; 50 % to 85.7 %; and 100 % to 88.8 % – a 4.9‑point (5.8 %) absolute gain over baseline. Under the same compute budget, RoboCasa365 reached only 57.5 %, 31.3 points lower than AXIS‑100 %.
Perturbation Analysis
AXIS‑100 % improves robustness to:
Sensor noise: +13.7 pp
Camera view changes: +11.3 pp
Initial robot pose: +3.8 pp
Background changes: +3.7 pp
Object layout changes: +2.6 pp
Small degradations appear under lighting (‑1.7 pp) and language perturbations (‑1.3 pp), indicating under‑covered dimensions.
Conclusions and Future Work
AXIS functions as a sustainable data engine that continuously generates, validates, augments, and evaluates robot manipulation data. The closed‑loop integration of task generation, asset deployment, web teleoperation, community collection, cleaning, Isaac Sim enhancement, and versioned publishing yields measurable model gains that scale with data size.
Planned extensions include additional robot bodies, sensor modalities, long‑horizon tasks, and a feedback loop where model failures directly guide task generation and data collection, moving from “continuous growth” to “targeted growth” based on capability gaps.
Paper: https://arxiv.org/abs/2607.21588
Project page: https://axisaiorg.github.io/AXIS-V1/
Code repository: https://github.com/AxisAIOrg/Axis-V1-Training
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
