8 Jargon Terms in Embodied Intelligence Data Collection: What Are Ego, GoPro, UMI, and Teleoperation?
The article demystifies eight common terms in embodied intelligence data collection—Ego, Exo, GoPro, UMI, Teleoperation, MoCap, Data Glove, and Robot Trajectory—explaining their meanings, how they differ, and why each is crucial for building effective robot learning datasets.
Ego: First‑Person Perspective Data
Ego (short for egocentric) refers to data captured from the viewpoint of the agent itself, typically by mounting a camera on the head and recording everyday actions such as drinking a cup or arranging items. It captures "what the robot sees" and is valuable because it records how an intelligent agent perceives the world from its own perspective.
Exo: Third‑Person Perspective
Exo is the opposite of Ego, providing an external view of the scene. A stationary camera records the whole body, the environment, and objects, offering a complementary view that fills in details Ego may miss, such as limb positions or full-body motion.
GoPro: The Capture Device, Not the Data Type
GoPro is often mistakenly called a data type. In reality, it is a brand of action cameras used to record Ego video. The distinction is that GoPro describes "what device is used" (the camera), while Ego describes "how the data is captured" (the perspective).
UMI (Universal Manipulation Interface)
UMI enables a human to demonstrate tasks using a robot‑like end‑effector (e.g., a gripper) instead of their bare hands. By performing actions with a device that mimics a robot’s manipulator, the resulting data aligns more closely with the robot’s kinematics, making it easier to map human demonstrations to robot motions.
Teleoperation: Human‑in‑the‑Loop Control
Teleoperation lets a human directly control a robot via a joystick, VR controller, or teach pendant. The system records the robot’s camera view, joint angles, end‑effector pose, commands, timestamps, and optionally force or tactile feedback. This creates paired data of "what the robot saw" and "what the human instructed it to do," which can later be used for imitation learning.
MoCap (Motion Capture)
MoCap captures the precise movement of a human body—joint angles, limb trajectories, and poses—rather than just video. It is essential for full‑body robots or humanoid agents because it provides detailed kinematic information that video alone cannot convey.
Data Glove
A data glove records hand articulation, finger bends, palm orientation, and sometimes force feedback. Since hand dexterity is critical for many manipulation tasks, glove data complements MoCap by providing high‑resolution information about hand motions that video cannot capture.
Robot Trajectory
Robot trajectory data records the actual states and actions of a robot during a task, including start pose, intermediate actions, environmental changes, and final outcome. High‑quality trajectories answer three questions: the environment state, the robot’s actions, and the resulting effect, making them the most valuable asset for training autonomous policies.
By understanding these four layers—viewpoint (Ego/Exo), capture device (GoPro, depth cameras), teaching method (UMI, Teleoperation, MoCap, Data Glove), and final data product (video, pose, joint states, trajectory)—practitioners can quickly assess project requirements, costs, and feasibility without getting lost in jargon.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
