Can General-Purpose Robots Arrive Soon? Inside the RoboDojo “Embodied Everest” Benchmark
The RoboDojo benchmark evaluates 30 robot manipulation policies across 42 simulated and 18 real-world tasks, revealing that the best models achieve only 8.8% success in simulation and 12.8% in reality, far behind human experts, and highlights gaps in generalization, memory, precision, long‑horizon execution, and open‑semantic understanding.
In the past year, keywords such as VLA, robot foundation models, and world models have dominated the embodied intelligence field, with demos showing robots that can stack bowls, insert tubes, pour water, and tidy tables, suggesting a shift from merely understanding language to actually performing tasks.
The core question remains: where do these models truly excel—only in simulation, or also in the real world, and can they complete entire task sequences reliably?
The team that previously released the RoboTwin series introduced RoboDojo, a unified simulation‑and‑real benchmark comprising 42 simulated tasks, 18 real‑world robot tasks, and 30 representative robot manipulation policies evaluated under a single standard.
Evaluation results show that the best simulation policy attains an average success rate of 8.80%, while the top real‑world policy reaches 12.8%; by contrast, human experts achieve 76.03% in simulation and 100% in reality, indicating a substantial performance gap.
RoboDojo assesses five ability dimensions: generalization (adapting to new backgrounds, lighting, objects, and clutter), memory (recalling previously seen information), precise manipulation (high‑accuracy insertion, alignment, contact), long‑horizon execution (completing multiple inter‑dependent steps), and open‑semantic understanding (interpreting unseen language instructions).
The real‑world suite covers three dual‑arm platforms (ARX5, Piper, Piper X), each with six tasks such as food preparation, tube insertion, charger insertion, bowl stacking, cup hanging, block cleaning, object classification, and backpack packing. Real‑world challenges include camera noise, calibration error, arm latency, gripper slip, unstable contact, and pose offsets, which can cause a strategy that is stable in simulation to degrade on hardware.
To ensure comparable and reproducible real‑world evaluation, RoboDojo‑RealEval standardizes hardware configuration, workspace layout, lighting, scene‑reset procedures, evaluation protocol, and a blind‑scoring system that judges both final outcomes and intermediate step completion.
On the simulation leaderboard, Hy‑Embodied‑0.5‑VLA leads with an average score of 13.07 and 8.80% success, still far behind the human benchmark of 76.03%. In the real‑world leaderboard, π0.5 ranks highest with a 12.8% success rate and an average score of 22.9, while other top models include InternVLA‑A1, GalaxeaVLA, Xiaomi‑Robotics‑0, and X‑VLA.
Open‑semantic tasks remain especially challenging; the best model achieves only 1.67% success, underscoring the difficulty of interpreting unseen instructions.
The analysis shows uneven progress: some models excel at visual recognition, others at motion execution or long‑horizon planning, but a truly general robot must perform stably across all dimensions.
Beyond the benchmark itself, RoboDojo provides two key infrastructure components: heterogeneous parallel simulation, which runs different tasks, objects, and layouts concurrently to boost evaluation efficiency, and XPolicyLab, a unified interface that normalizes observation‑action formats, preprocessing, training scripts, and deployment environments for diverse policies.
Future plans aim to extend RoboDojo to dexterous manipulation, full‑body humanoid tasks, tactile manipulation, and mobile operations, positioning it as a sustainable arena for embodied intelligence research.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
