GigaBrain-0.7 Unveiled: First Open‑Source Embodied Model with Ground‑Breaking System‑3 Dual‑Tower Architecture
GigaBrain-0.7, the latest embodied foundation model released by JiJia Vision, introduces the novel System‑3 paradigm that embeds a world model into the robot's decision loop, leverages a dual "data pyramid + algorithm pyramid" framework, and achieves top rankings across all RoboColiseum sub‑benchmarks, demonstrating unprecedented real‑world performance and scaling behavior.
At the 2026 World Robot Conference, JiJia Vision showcased GigaBrain-0.7, a new generation embodied foundation model that directly topped the RoboColiseum leaderboard by securing first place in all four capability sub‑rankings within five days of release.
System‑3 Concept : GigaBrain-0.7 is the first model to integrate a world model into the robot's real‑time decision loop, enabling the robot to "pre‑simulate judgment before executing actions." This is achieved through the System‑3 architecture, which generates future action‑value feedback and visual predictions that are fed back into the model's prompt.
Three‑Layer Algorithm Pyramid :
Layer 1 – World Simulation : Uses a large‑scale visual‑language model (VLM) as the backbone, enhanced with a Temporal‑Spatial Block and relative position/time encodings to solve long‑range temporal reasoning.
Layer 2 – Action Alignment : Built on a MoT architecture with Flow Matching, it shares most Transformer parameters across different robot morphologies via a Robot ID routing mechanism. A Soft Knowledge Insulation layer protects the VLM’s generalization while boosting task‑specific precision.
Layer 3 – Experience Reinforcement : Combines supervised fine‑tuning, offline RL (using rollout data and human‑in‑the‑loop annotations), and online continuous RL to progressively improve success rates on high‑difficulty, long‑tail tasks.
Dual‑Pyramid Data Strategy : The model is trained on a massive multimodal dataset comprising nearly 3 billion image‑language pairs, tens of thousands of hours of video, and over 11 200 hours of high‑quality real‑robot data collected with self‑designed UMI (hand‑held) and EGO (first‑person) sensors. An additional 5 000 hours of generated high‑fidelity simulation data further enriches the training corpus.
Scaling Validation : Experiments show that increasing data volume mitigates over‑fitting and improves generalization, while adding UMI/EGO data directly boosts real‑world success rates. Progressive pre‑training stages demonstrate a monotonic rise in real‑robot task success, eventually achieving 100 % success across 24 evaluated tasks without task‑specific fine‑tuning.
Benchmark Performance : Compared against leading open‑source and proprietary models on the RoboColiseum platform (both simulation and real‑robot settings on AgileX PiPER and Maker H01), GigaBrain-0.7 consistently outperforms competitors, especially on the Maker H01 platform where it exhibits a "gap‑leading" advantage. It also completes complex long‑duration tasks (over 20 minutes and more than 10 tasks) in a single uninterrupted run.
Resources : Project page – https://gigaai.cc/blog/gigabrain07; Technical report – https://arxiv.org/abs/2608.15875; Code repository (coming soon) – https://github.com/open-gigaai/giga-brain-0; Model hub (coming soon) – https://huggingface.co/open-gigaai/models.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
