Berkeley PhD Thesis: Building Generalist Robots via Data, Policy, Cross-Embodiment & Memory
Berkeley PhD thesis presents a four-part framework for generalist robots: BridgeData V2 dataset (60k trajectories), Octo policy (800k demos, 29% better zero-shot), CrossFormer (single policy across 20 embodiments, 73% success), and MEM (multi-scale memory for 15-min tasks).
Paper Reference
Title: Building Generally Intelligent Robots
Author: Homer Walke
Institution: University of California, Berkeley
Technical Report: UCB/EECS-2026-267
Date: August 14, 2026
Paper Link: https://www2.eecs.berkeley.edu/Pubs/TechRpts/2026/EECS-2026-267.pdf
Introduction: The Challenge of Robot Foundation Models
The thesis addresses a core question: if language and vision have achieved foundation-model generalization via massive data and unified models, can robotics follow a similar path to train agents that generalize across tasks, environments, and robot bodies? The author argues that simply transferring large models to robots is insufficient; instead, the problem must be decomposed into four layers: (1) data covering diverse tasks, objects, and environments; (2) a policy architecture handling language, goal images, multi-view observations, and continuous actions; (3) cross-embodiment support for single-arm, dual-arm, mobile, and quadruped robots; (4) scalable memory for long-horizon tasks requiring context beyond seconds.
BridgeData V2: Data Coverage as the First Bottleneck
Robot generalization research is hindered by small datasets and evaluations too close to training distributions. BridgeData V2 extends the original Bridge Dataset to 60,096 trajectories (50,365 expert demos + 9,731 scripted), covering 13 skill types, 24 environments, and 100+ objects — 7× the original scale. The dataset is explicitly organized around generalization variables: skill categories (object manipulation, cloth, block stacking, granular sweeping, mixed object/environment) and environment types (toy kitchen, tabletop, toy sink). This design enables systematic evaluation on unseen objects, unseen environments, and their combination.
The dataset supports multiple offline learning paradigms: goal-conditioned behavioral cloning (using future observations or goal images), language-conditioned behavioral cloning (natural language instructions), and offline reinforcement learning. Experiments show that larger models, larger training sets, and higher data diversity all improve generalization, especially on unseen objects and environments. Positive transfer across related manipulation skills is observed: joint training on diverse skills outperforms narrow-distribution training. The key takeaway: a generalist robot policy cannot emerge from architecture alone; without sufficiently diverse data, models learn narrow shortcuts regardless of capacity.
Octo: A Generalist Robot Policy
Built on the data foundation, Octo is an open-source generalist policy designed as a fine-tunable foundation model for robotics. It uses a Transformer backbone that tokenizes task descriptions (language or goal images) and robot observations (multi-camera views, proprioception), then outputs continuous actions via a diffusion action head — emphasized as critical for modeling continuous robot actions.
Octo is pre-trained on ~800k demonstrations from 25 datasets within Open X-Embodiment, offering broader manipulation coverage than earlier generalist policies. Two checkpoints are released: Octo-Small (27M parameters) and Octo-Base (93M parameters), sized for real-time control rather than matching massive vision-language models. A key design choice is swappable input/output interfaces: new cameras, proprioception, or action spaces unseen during pre-training can be integrated at fine-tuning via new tokenizers or action heads without rewriting the entire policy.
Evaluated across 4 institutions and 9 real robot setups, Octo achieves strong zero-shot control on in-distribution tasks. Under language conditioning, it outperforms open-source RT-1-X by 29% average success rate and matches the larger RT-2-X on WidowX and RT-1 Robot tasks. More importantly, fine-tuning Octo on new tasks and robots yields a 52% average improvement over the next-best baseline across 6 evaluation settings, demonstrating that large-scale pre-training primarily moves policy parameters to a region more adaptable to new scenarios. Ablation studies confirm that vision Transformer backbone, diverse training data, diffusion action head, model scale, and flexible I/O interfaces all contribute critically. Limitations acknowledged: focus remains on single/dual-arm manipulation; zero-shot degrades on out-of-distribution skills; wrist-camera utilization needs improvement.
CrossFormer: Single Policy Across Multiple Embodiments
If a generalist policy only controls one arm type, it falls short of "general." Real robots vary widely in sensors, action dimensions, control frequencies, and task objectives. CrossFormer extends the Octo-style architecture to handle this heterogeneity by allowing each embodiment its own observation tokens and action heads while maximizing shared Transformer backbone parameters. This preserves cross-robot knowledge sharing without forcing incompatible action spaces into a single fixed vector.
Training uses a cross-embodiment data mix from Open X-Embodiment covering 20 robot bodies: single-arm manipulators, dual-arm robots, navigation platforms, and quadrupeds. The goal is not to make all robots execute the same action, but to let one policy framework complete each robot's respective tasks via its specific I/O interfaces. Real-world evaluation shows CrossFormer achieves 73% average success rate, compared to 67% for a same-architecture single-robot baseline and 51% for the best prior method. This indicates cross-embodiment joint training causes no significant negative transfer and matches or exceeds specialist policies in multiple settings. The conclusion is measured: CrossFormer proves feasibility of a unified policy across embodiments and shows early positive transfer signals, but cannot yet claim strong positive transfer across all bodies. It crosses an important threshold — one model can structurally accommodate multiple robots and remain competitive on real tasks — a prerequisite for future robot foundation platforms that must absorb data from diverse hardware rather than retraining isolated models per robot.
MEM: Adding Multi-Scale Embodied Memory
Prior policies operate on short observation windows, yet many real tasks require long-term context: open drawer → place object inside; locate tool based on position seen minutes ago; adjust strategy after a failed grasp. Feeding all history frames into a large model is computationally infeasible for high-dimensional video. MEM equips vision-language-action models with long-horizon memory while preserving real-time control via a multi-scale design: long-term semantic events are compressed into language memory; short-term dense visual changes are compressed by a video encoder.
The policy splits into high-level and low-level components. The high-level policy updates language memory, summarizing key past events into readable text (completed subtasks, object locations, prior success/failure). The low-level policy conditions on current goal, short-term video memory, and high-level sub-task instructions to output continuous actions. This avoids stuffing minutes of video into context: language memory retains long-term semantic state, video encoder retains fine-grained details for near-term manipulation. Experiments use observation memory up to 18 frames (54 seconds) while tasks scale to 15 minutes.
Long-horizon tasks include kitchen tidying, recipe-based item preparation, and space cleanup. Results: strong baselines without memory struggle; language-only or video-only memory are insufficient; full MEM significantly improves task progress, demonstrating complementarity of semantic long-term and visual short-term memory. MEM also shows in-context adaptation: robots adjust grasp height, opening direction, or subsequent manipulation based on prior failures or environment changes — critical for real deployment where static, single-success environments are unrealistic. MEM addresses the "can it sustain a long task" dimension, completing the progression from "can it do it" (Octo) and "can it do it on another body" (CrossFormer) to "can it keep doing it over time."
Conclusion and Open Problems
The thesis outlines a four-step path: (1) define generalization with sufficiently rich data (BridgeData V2); (2) absorb multi-source demonstrations with a scalable Transformer policy (Octo); (3) unify control across embodiments (CrossFormer); (4) extend from short-horizon control to long-horizon tasks via multi-scale memory (MEM). These are not independent projects but consecutive rungs toward generally intelligent robots.
Open problems remain: (1) robot data scale still lags far behind language/vision; low-cost, high-quality, continuous collection of diverse interaction data is a core bottleneck. (2) Cross-embodiment transfer works but lacks systematic theory/evaluation for when positive vs. negative transfer occurs. (3) Long-term memory relies on structured semantic compression; maintaining stable, correctable, auditable memory in open environments is unsolved. (4) Real deployment demands safety, robustness, anomaly recovery, and human-robot collaboration — challenges unlikely to be solved by offline imitation learning alone. The thesis's value lies in decomposing "generally intelligent robots" from an abstract vision into an experimentally grounded, reproducible, extensible engineering path, showing robot foundation models have taken solid steps in data scale, model structure, and real evaluation, while reminding us that robot intelligence's difficulty lies not just in understanding the world but in continuously, reliably, and low-latency changing it.
Code example
来源:专知
本文
约5000字
,建议阅读
10
分钟
这篇论文的价值在于把“通用智能机器人”从抽象愿景拆成可实验、可复现、可扩展的工程路径。Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data Party THU
Official platform of Tsinghua Big Data Research Center, sharing the team's latest research, teaching updates, and big data news.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
