Tagged articles

embodied AI

163 articles · Page 2 of 2
Xiaomi Tech
Xiaomi Tech
Apr 27, 2026 · Artificial Intelligence

Xiaomi‑Robotics‑0: 20‑Hour Post‑Training Enables Seamless Earphone‑Box Assembly (Open‑Source)

The article details how Xiaomi‑Robotics‑0 achieves precise earphone‑to‑case insertion after only 20 hours of post‑training, outlines the sub‑millimetre precision challenges, presents a triple‑strategy (asynchronous execution, adaptive loss re‑weighting, Λ‑shape attention mask and random masking) to avoid the "lazy effect", and releases the full pipeline and code as open source for the robotics community.

Asynchronous ExecutionXiaomi Roboticsaction prefixing
0 likes · 6 min read
Xiaomi‑Robotics‑0: 20‑Hour Post‑Training Enables Seamless Earphone‑Box Assembly (Open‑Source)
Meituan Technology Team
Meituan Technology Team
Apr 23, 2026 · Artificial Intelligence

LARYBench Introduces an ImageNet‑Style Benchmark for Embodied Action Representations Learned from Human Video

LARYBench (Latent Action Representation Yielding Benchmark) provides the first systematic, ImageNet‑scale evaluation for implicit action representations derived from large‑scale human video, decoupling representation quality from downstream control, and shows that general‑purpose vision models outperform specialized embodied models in both action generalization and control precision across diverse robot morphologies and environments.

Vision-Language-Actionaction representationbenchmark
0 likes · 13 min read
LARYBench Introduces an ImageNet‑Style Benchmark for Embodied Action Representations Learned from Human Video
HyperAI Super Neural
HyperAI Super Neural
Apr 23, 2026 · Artificial Intelligence

Task Tokens Cut Per-Task Trainable Parameters 125× and Boost Convergence 6× for Embodied AI

The Task Tokens method introduced by an Israeli research team reduces the number of trainable parameters per task by up to 125‑fold and speeds up convergence by six times, while preserving the flexibility of Behavior Foundation Models and demonstrating strong performance, robustness, and compatibility across a suite of embodied control tasks.

Behavior Foundation ModelsMulti-Modal PromptingPPO
0 likes · 13 min read
Task Tokens Cut Per-Task Trainable Parameters 125× and Boost Convergence 6× for Embodied AI
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Apr 22, 2026 · Artificial Intelligence

How to Build an End‑to‑End Hand‑Video to VLA Data Pipeline on Alibaba Cloud PAI with Data‑Juicer

This article details a step‑by‑step, distributed pipeline built on Alibaba Cloud PAI using Data‑Juicer and Ray that transforms raw egocentric hand videos into LeRobot v2.0‑compatible Vision‑Language‑Action (VLA) training data, covering video splitting, frame extraction, camera calibration, 3D hand reconstruction, pose estimation, action captioning, and export, with code snippets, performance numbers, and references.

Data PipelineData-JuicerDistributed Computing
0 likes · 29 min read
How to Build an End‑to‑End Hand‑Video to VLA Data Pipeline on Alibaba Cloud PAI with Data‑Juicer
Code Mala Tang
Code Mala Tang
Apr 22, 2026 · Artificial Intelligence

How LeWorldModel Achieves Stable End‑to‑End World Modeling with Just Two Losses

LeWorldModel, a 2026 JEPA‑based world model introduced by Yann LeCun and collaborators, solves representation collapse with a minimalist two‑loss objective, delivering a 15‑million‑parameter system that trains in hours, runs 48× faster than prior baselines, and reaches near‑SOTA performance on robot control benchmarks.

JEPAdeep learningembodied AI
0 likes · 6 min read
How LeWorldModel Achieves Stable End‑to‑End World Modeling with Just Two Losses
Architect's Must-Have
Architect's Must-Have
Apr 21, 2026 · Artificial Intelligence

30 Essential AI Agent Concepts: From LLMs to Multi‑Agent Systems

This comprehensive guide systematically explains thirty core terms of AI agents—covering foundational large language models, fine‑tuning techniques, multimodal vision‑language models, agent architectures such as ReAct and CoT, tool‑calling protocols, retrieval‑augmented generation, workflow orchestration, and emerging product forms like autonomous and embodied agents—while detailing the reasoning, trade‑offs, and concrete examples that shape modern agent engineering.

AI agentsLarge Language ModelsPrompt Engineering
0 likes · 36 min read
30 Essential AI Agent Concepts: From LLMs to Multi‑Agent Systems
Machine Heart
Machine Heart
Apr 20, 2026 · Industry Insights

The Toughest Dexterous Robotic Hand Yet: OmniHand 3 Ultra‑T, Lite, and OmniPicker 3 Unveiled

At the 2024 ZhiYuan Partner Conference, the company introduced three new rope‑driven dexterous hands—OmniHand 3 Ultra‑T, OmniHand 3 Lite, and OmniPicker 3—detailing their technical routes, performance specs, ruggedness improvements, and open‑source ecosystem that aim to make high‑precision manipulation affordable and reliable for research and industry.

OmniHandOmniPickerdexterous hand
0 likes · 18 min read
The Toughest Dexterous Robotic Hand Yet: OmniHand 3 Ultra‑T, Lite, and OmniPicker 3 Unveiled
Machine Heart
Machine Heart
Apr 20, 2026 · Artificial Intelligence

Deployment Era Starts: How One Firm Delivered Seven Turnkey Embodied‑AI Solutions Without Selling Robots

ZhiYuan announced four new robot bodies, six AI models and seven standardized productivity solutions, backed by a full‑stack AIMA ecosystem and a massive data network, achieving 10,000 mass‑produced robots by 2026, 39% market share in 2025 and revenue surpassing 1 billion yuan, marking the first year of the embodied‑AI deployment era.

AI modelsEcosystemdeployment
0 likes · 14 min read
Deployment Era Starts: How One Firm Delivered Seven Turnkey Embodied‑AI Solutions Without Selling Robots
Architect's Must-Have
Architect's Must-Have
Apr 20, 2026 · Industry Insights

How Humanoid Robots Beat the Human Marathon Record – Inside the 2026 Beijing Race

The 2026 Beijing Yizhuang half‑marathon saw over 300 humanoid robots compete, with the champion "Lightning" finishing in 50 minutes 26 seconds—three times faster than the previous year and faster than the human world record—while the event revealed six core technical breakthroughs, a rapid rise in autonomous navigation, a dominant Chinese supply chain, and a roadmap for future industrial and consumer applications.

autonomous navigationembodied AIhumanoid robots
0 likes · 22 min read
How Humanoid Robots Beat the Human Marathon Record – Inside the 2026 Beijing Race
Machine Heart
Machine Heart
Apr 19, 2026 · Artificial Intelligence

Gaode’s Fully Autonomous Embodied Robot Conquers Guide‑Blind Challenge at Yizhuang Marathon

Gaode’s four‑legged robot "Gaode Tutu" demonstrated fully autonomous navigation and manipulation in an open‑world marathon, tackling the guide‑blind task with a visually impaired teen and achieving state‑of‑the‑art results on multiple navigation and manipulation benchmarks using its ABot full‑stack system.

ABotNavigationembodied AI
0 likes · 19 min read
Gaode’s Fully Autonomous Embodied Robot Conquers Guide‑Blind Challenge at Yizhuang Marathon
Amap Tech
Amap Tech
Apr 19, 2026 · Artificial Intelligence

From Pixels to the Physical World: Inside Gaode’s ABot Full‑Stack

Gaode leverages 20 years of spatiotemporal data to unveil a three‑layer ABot stack—World Model, navigation (N series), manipulation (M series) and a Harness architecture—that embeds physical laws, self‑evolves through a dual data‑training engine, and achieves benchmark‑leading performance across wheeled, quadruped and humanoid robots.

Navigationbenchmarkembodied AI
0 likes · 16 min read
From Pixels to the Physical World: Inside Gaode’s ABot Full‑Stack
Machine Heart
Machine Heart
Apr 18, 2026 · Artificial Intelligence

Why Embodied Data Is the Biggest Gold Mine: Inside the World’s First Hundred‑Billion‑Scale Multimodal Data Cloud Mall

Paxini, together with JD Cloud, Tencent Cloud, and Baidu Intelligent Cloud, launches the world’s first hundred‑billion‑scale, full‑modal, high‑degree‑of‑freedom embodied AI data cloud mall, offering instant online data procurement, end‑to‑end model training pipelines, and validated performance gains in both lab and real‑world robot tasks.

Large-Scale DataMultimodal Datadata cloud marketplace
0 likes · 13 min read
Why Embodied Data Is the Biggest Gold Mine: Inside the World’s First Hundred‑Billion‑Scale Multimodal Data Cloud Mall
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Apr 17, 2026 · Artificial Intelligence

LARYBench: An ImageNet‑Scale Benchmark Unlocks Embodied AI Generalization

Researchers introduce LARYBench, the first large‑scale benchmark for evaluating implicit action representations in embodied AI, providing over 1.2 million annotated video clips, a unified metric for motion semantics, and extensive experiments showing that general visual encoders outperform specialized robot models in action understanding and control.

LARYBenchVision Encodersaction representation
0 likes · 12 min read
LARYBench: An ImageNet‑Scale Benchmark Unlocks Embodied AI Generalization
Machine Heart
Machine Heart
Apr 14, 2026 · Artificial Intelligence

Why Binary Success Rate Is Obsolete: Introducing PRM-as-a-Judge for Dense Evaluation of Embodied Tasks

The article critiques binary success rate for long‑horizon robotic tasks, proposes the PRM-as-a-Judge framework with a potential‑based progress signal and the three‑layer OPD metric suite, validates it on the RoboPulse benchmark, and shows how it yields fine‑grained, diagnostic insights into policy performance.

OPDRoboPulsedense metrics
0 likes · 20 min read
Why Binary Success Rate Is Obsolete: Introducing PRM-as-a-Judge for Dense Evaluation of Embodied Tasks
Machine Heart
Machine Heart
Apr 13, 2026 · Artificial Intelligence

How Six‑Dimensional Force Data Powers China’s First Full‑Perception VTLA Model

The article analyzes how Kepler Robotics’ dual‑path, six‑degree‑of‑freedom force‑tactile data collection system overcomes the scaling bottleneck of embodied AI, enabling a VTLA model that integrates vision, language, action and tactile feedback to achieve near‑perfect industrial assembly performance.

Data CollectionKepler RoboticsVTLA model
0 likes · 14 min read
How Six‑Dimensional Force Data Powers China’s First Full‑Perception VTLA Model
Machine Heart
Machine Heart
Apr 11, 2026 · Artificial Intelligence

How 100,000 Hours of Human Data Propelled Psi‑R2 to Lead MolmoSpaces

Lingchu AI demonstrates that scaling human‑operation data to nearly 100,000 hours, combined with a two‑model system and reinforcement learning, can replace costly robot‑teleoperation data and achieve top performance on the MolmoSpaces benchmark.

Psi-R2Psi-W0embodied AI
0 likes · 12 min read
How 100,000 Hours of Human Data Propelled Psi‑R2 to Lead MolmoSpaces
Machine Heart
Machine Heart
Apr 10, 2026 · Artificial Intelligence

Why Generalist’s Success Shifts Embodied AI Competition From Models to Infrastructure

The launch of Generalist AI’s GEN‑1 model demonstrates a breakthrough in success rate, speed and resilience, but the article argues that the true competitive frontier has moved from model performance to the underlying data, simulation and evaluation infrastructure that enables continuous learning and scalable testing for embodied intelligence.

AI modelsData InfrastructureSimulation
0 likes · 12 min read
Why Generalist’s Success Shifts Embodied AI Competition From Models to Infrastructure
Machine Heart
Machine Heart
Apr 10, 2026 · Artificial Intelligence

How a Chinese Company Swept the Embodied Intelligence Olympics with Faster, Precise, Low‑Data Robotics

A Chinese robotics firm leveraged a self‑developed VLA model to win all three core tasks at Benjie’s Embodied Intelligence Olympics—peeling oranges, unlocking doors, and flipping socks—outperforming the industry leader Physical Intelligence by up to 35% faster speed, using 30% fewer samples and achieving higher precision in real‑world, fully autonomous scenarios.

VLA modelbenchmark competitionembodied AI
0 likes · 16 min read
How a Chinese Company Swept the Embodied Intelligence Olympics with Faster, Precise, Low‑Data Robotics
Machine Heart
Machine Heart
Apr 7, 2026 · Artificial Intelligence

A Comprehensive Survey of Tactile‑Based Multimodal Fusion in Embodied Intelligence

This survey reviews state‑of‑the‑art research up to Q1 2026 on integrating tactile sensing with vision and language for embodied AI, presenting a four‑stage fusion pipeline, a hierarchical taxonomy of datasets, methods, sensors, and highlighting current evaluation challenges and future directions.

Multimodal Fusiondatasetsembodied AI
0 likes · 13 min read
A Comprehensive Survey of Tactile‑Based Multimodal Fusion in Embodied Intelligence
Machine Heart
Machine Heart
Apr 7, 2026 · Artificial Intelligence

How Qianxun Raised ¥3 B in 30 Days: AI‑Powered Robotics Secrets

Qianxun Intelligent secured ¥30 billion in funding within a month, leveraged a scaling‑law data engine and the Spirit v1.5 VLA model to achieve breakthrough robot performance, and demonstrated the commercial loop through deployments at JD.com retail and CATL battery lines.

Data CollectionQianxun Intelligentembodied AI
0 likes · 12 min read
How Qianxun Raised ¥3 B in 30 Days: AI‑Powered Robotics Secrets
Machine Heart
Machine Heart
Apr 3, 2026 · Artificial Intelligence

Manifold AI’s WorldScape Tops WorldScore, Outperforming Li Fei‑Fei’s Team

Manifold AI’s WorldScape model claimed the top spot on the WorldScore benchmark, beating leading labs such as Li Fei‑Fei’s team, MIT, Alibaba and Runway, while using an order‑of‑magnitude fewer parameters, integrating generation and control, delivering real‑time 6‑16 FPS interactive 3‑D output with stable geometry and world‑state memory.

Manifold AIWorldScapeWorldScore
0 likes · 9 min read
Manifold AI’s WorldScape Tops WorldScore, Outperforming Li Fei‑Fei’s Team
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Mar 31, 2026 · Artificial Intelligence

GigaWorld-1 Tops WorldArena Benchmark, Surpassing Google and Nvidia

GigaWorld-1, the latest embodied world model from Jiji Vision, clinched the global #1 spot on the WorldArena benchmark—beating Google, Nvidia, and Alibaba—with a comprehensive score over 60, excelling in physics adherence (+16%), near‑perfect 3D accuracy, and leading visual quality, while leveraging explicit action modeling, a differentiable physics engine, massive robot video data, and open‑source releases that have already attracted over 16,000 downloads.

Open Sourcebenchmarkembodied AI
0 likes · 7 min read
GigaWorld-1 Tops WorldArena Benchmark, Surpassing Google and Nvidia
Amap Tech
Amap Tech
Mar 30, 2026 · Artificial Intelligence

ABot-M0: A Unified VLA Framework Solving the One‑Brain Many‑Forms Robotics Challenge

ABot-M0 is an open‑source Vision‑Language‑Action foundation model that unifies fragmented robot data, introduces Action Manifold Learning for smoother action prediction, and offers a plug‑and‑play dual‑stream perception architecture, achieving state‑of‑the‑art results on major manipulation benchmarks.

Foundation Modelaction manifold learningembodied AI
0 likes · 4 min read
ABot-M0: A Unified VLA Framework Solving the One‑Brain Many‑Forms Robotics Challenge
Old Meng AI Explorer
Old Meng AI Explorer
Mar 30, 2026 · Industry Insights

Why SoftBank’s $40B Bet Signals a New Era of AI Competition

The article analyzes SoftBank’s $40 billion unsecured loan to double‑down on OpenAI, the launch of OpenAI’s GPT‑5.4 with million‑token context, Google’s Gemini 3.1 Flash Live voice model, Chinese AI’s market surge, the rise of embodied intelligence, AI agents becoming autonomous coworkers, and the broader industry polarization between massive funding and job displacement, offering a comprehensive snapshot of AI’s 2026 landscape.

AIIndustry TrendsOpenAI
0 likes · 22 min read
Why SoftBank’s $40B Bet Signals a New Era of AI Competition
Old Meng AI Explorer
Old Meng AI Explorer
Mar 26, 2026 · Industry Insights

How AI Shifted From Chatbots to Digital Employees in March 2026

In March 2026, breakthrough models like GPT‑5.4 and Claude 4.6 introduced native computer control and million‑token contexts, Chinese video AI topped global rankings, capital poured over ¥200 billion into embodied intelligence, and AI agents began scaling from tools to digital employees across enterprises.

AIAI video generationcapital investment
0 likes · 25 min read
How AI Shifted From Chatbots to Digital Employees in March 2026
HyperAI Super Neural
HyperAI Super Neural
Mar 25, 2026 · Artificial Intelligence

Low‑Barrier Deployment of NVIDIA’s Latest Physical AI Models for Humanoid Robots, Motion Generation, and Diffusion Fine‑Tuning

The article introduces NVIDIA’s Physical AI suite announced at GTC 2026—including Isaac GR00T, SOMA‑X, Kimodo, and FDFO—explains each model’s architecture and purpose, and provides one‑click online tutorials that let developers experiment with humanoid robotics, human‑body modeling, motion generation, and diffusion model fine‑tuning at minimal cost.

FDFOIsaac GR00TKimodo
0 likes · 8 min read
Low‑Barrier Deployment of NVIDIA’s Latest Physical AI Models for Humanoid Robots, Motion Generation, and Diffusion Fine‑Tuning
Amap Tech
Amap Tech
Mar 20, 2026 · Artificial Intelligence

How ABot-PhysWorld Achieves Physical Consistency in Embodied Video Generation

ABot-PhysWorld introduces a physically consistent video generation framework for embodied AI, leveraging the PAI‑Bench benchmark, large‑scale multi‑modal data, DPO preference alignment, and dense action maps to surpass SOTA models in both visual quality and physical plausibility across diverse robotic tasks.

Physical ConsistencyVideo Generationbenchmark
0 likes · 15 min read
How ABot-PhysWorld Achieves Physical Consistency in Embodied Video Generation
AI Explorer
AI Explorer
Mar 17, 2026 · Artificial Intelligence

RISE Enables Breakthrough in Vision‑Language‑Action Learning for Embodied AI

The article examines the limitations of vision‑language‑action (VLA) models in real‑world tasks, explains how the RISE technique from Hong Kong University uses internal simulation, reflection and imagination to cut training costs by an order of magnitude, and discusses its implications for future embodied AI.

RISEVLAembodied AI
0 likes · 6 min read
RISE Enables Breakthrough in Vision‑Language‑Action Learning for Embodied AI
HyperAI Super Neural
HyperAI Super Neural
Mar 9, 2026 · Artificial Intelligence

Physics‑Informed GNN Breakthrough for Accurate, Real‑Time Multi‑Body Dynamics

Researchers from EPFL introduce DYNAMI‑CAL GraphNet, a graph neural network that embeds linear and angular momentum conservation, delivering highly accurate, interpretable and real‑time predictions for complex multi‑body systems across robotics, aerospace and materials science, and outperforming existing baselines on four diverse benchmark datasets.

DYNAMI‑CAL GraphNetSimulationembodied AI
0 likes · 16 min read
Physics‑Informed GNN Breakthrough for Accurate, Real‑Time Multi‑Body Dynamics
AI Frontier Lectures
AI Frontier Lectures
Mar 5, 2026 · Artificial Intelligence

Can Robots Navigate Unseen Spaces with Only Language? EvoNav’s Zero‑Shot Vision‑Language Breakthrough

The EvoNav framework from Nanjing University of Science and Technology tackles the last‑hundred‑meter challenge of embodied navigation by integrating a Future Chain‑of‑Thought and a Historical Experience chain, achieving significant zero‑shot performance gains on VLN‑CE benchmarks and real‑world robot tests, with code released on GitHub.

EvoNavFuture Chain of ThoughtHistorical Experience
0 likes · 6 min read
Can Robots Navigate Unseen Spaces with Only Language? EvoNav’s Zero‑Shot Vision‑Language Breakthrough
AI Explorer
AI Explorer
Feb 28, 2026 · Artificial Intelligence

How VLAW Unites World Models and Visual Language Models to Advance Embodied AI

The VLAW framework, developed by researchers from Tsinghua and Stanford, integrates high‑fidelity world models with visual‑language models, enabling real‑time physical interaction and intent understanding, which could dramatically improve training efficiency for embodied robots and mark a milestone toward safe, autonomous agents in complex real‑world environments.

SimulationVLAWWorld Models
0 likes · 6 min read
How VLAW Unites World Models and Visual Language Models to Advance Embodied AI
Sohu Tech Products
Sohu Tech Products
Feb 25, 2026 · Artificial Intelligence

How to Replicate the Spring Festival Robot Dance: A Complete Video‑to‑Robot Motion Guide

This tutorial walks you through building a full video‑to‑robot motion pipeline—from installing the necessary repositories and environments, configuring GMR and PromptHMR, running command‑line tools, launching a multilingual Web UI, to exporting multi‑person trajectories and MuJoCo simulations—while highlighting common pitfalls and advanced considerations.

GitHubSimulationTutorial
0 likes · 15 min read
How to Replicate the Spring Festival Robot Dance: A Complete Video‑to‑Robot Motion Guide
PaperAgent
PaperAgent
Feb 25, 2026 · Artificial Intelligence

How RynnBrain Unifies Perception, Reasoning, and Planning for Embodied AI

RynnBrain, an open‑source unified spatiotemporal foundation model from Alibaba DAMO Academy, integrates perception, localization, physics‑based reasoning and planning across 2 B, 8 B and 30 B MoE scales, handles multimodal visual inputs, and outperforms existing models on over 20 embodied benchmarks.

AlibabaFoundation ModelMultimodal
0 likes · 3 min read
How RynnBrain Unifies Perception, Reasoning, and Planning for Embodied AI
HyperAI Super Neural
HyperAI Super Neural
Feb 19, 2026 · Artificial Intelligence

World Model & VLA Breakthroughs: Top Papers from NVIDIA, ByteDance, Tsinghua and Others

This roundup highlights six recent embodied AI papers that advance world models and vision‑language‑action (VLA) techniques, covering DreamDojo's massive first‑person video model, LingBot‑World simulator, Agent World Model generator, BagelVLA, ACoT‑VLA, and the closed‑loop World‑VLA‑Loop framework.

Synthetic EnvironmentsVision-Language-ActionWorld Models
0 likes · 8 min read
World Model & VLA Breakthroughs: Top Papers from NVIDIA, ByteDance, Tsinghua and Others
HyperAI Super Neural
HyperAI Super Neural
Feb 14, 2026 · Artificial Intelligence

Beyond Visual Realism: WorldArena Benchmark Reveals the Capability Gap in Embodied World Models

WorldArena introduces a unified benchmark that evaluates generated videos not only for visual fidelity but also for embodied task functionality across six dimensions, exposing a stark gap between visual realism and practical usefulness and providing a composite EWMScore to compare models.

Physical ConsistencyVideo GenerationWorldArena
0 likes · 9 min read
Beyond Visual Realism: WorldArena Benchmark Reveals the Capability Gap in Embodied World Models
Amap Tech
Amap Tech
Feb 13, 2026 · Artificial Intelligence

How ABot‑M0 Achieves Generalist Robot Intelligence with Action Manifold Learning

ABot‑M0 tackles the three long‑standing "Babel Tower" challenges of embodied AI—data fragmentation, inconsistent representations, and training mismatches—by releasing the massive UniACT dataset, introducing Action Manifold Learning for direct action prediction, and designing a plug‑and‑play dual‑path perception architecture that outperforms prior models on multiple robot benchmarks.

action manifold learningdatasetembodied AI
0 likes · 14 min read
How ABot‑M0 Achieves Generalist Robot Intelligence with Action Manifold Learning
HyperAI Super Neural
HyperAI Super Neural
Feb 5, 2026 · Artificial Intelligence

16 Embodied AI Datasets Covering Grasping, QA, Logical and Trajectory Reasoning

This article compiles sixteen high‑quality embodied AI datasets—including simulation assets, robot motion retargeting, indoor scenes, multimodal benchmarks, grasping, question answering, trajectory reasoning and large‑scale robot learning collections—detailing their scope, size, and download links to support research on agents that perceive, decide, and act in the physical world.

MultimodalSimulationdataset
0 likes · 15 min read
16 Embodied AI Datasets Covering Grasping, QA, Logical and Trajectory Reasoning
HyperAI Super Neural
HyperAI Super Neural
Jan 23, 2026 · Artificial Intelligence

Embodied AI Resources: Datasets, Modeling, Papers (Nvidia, ByteDance, Xiaomi)

This article compiles a comprehensive set of embodied AI resources, including large‑scale robot learning datasets such as BC‑Z (32 GB) and DexGraspVLA (7 GB), interactive world‑modeling frameworks like HY‑World 1.5, open‑source LLM deployments, and recent research papers from Nvidia, ByteDance, Xiaomi and leading universities, each with download links and brief summaries.

AI research papersOpen-source ModelsWorld Modeling
0 likes · 14 min read
Embodied AI Resources: Datasets, Modeling, Papers (Nvidia, ByteDance, Xiaomi)
DataFunSummit
DataFunSummit
Jan 17, 2026 · Artificial Intelligence

How UnrealZoo Accelerates Embodied AI Research with High‑Fidelity Simulation

This article outlines the evolution from traditional AI to embodied intelligence, explains the Vision‑Language‑Action (VLA) paradigm, highlights data‑collection bottlenecks, introduces the UnrealZoo simulation platform built on Unreal Engine, and showcases real‑world case studies and future challenges for embodied AI research.

Data CollectionSimulationUnreal Engine
0 likes · 16 min read
How UnrealZoo Accelerates Embodied AI Research with High‑Fidelity Simulation
PaperAgent
PaperAgent
Jan 12, 2026 · Artificial Intelligence

How Mental World Models Are Redefining Embodied AI: A Comprehensive Review

This review introduces the Mental World Model (MWM) as a new cognitive layer for Embodied AI, compares it with traditional Physical World Models, outlines 19 Theory‑of‑Mind methods, 26 evaluation benchmarks, and discusses key challenges and future research directions.

Mental World ModelModel-BasedTheory of Mind
0 likes · 9 min read
How Mental World Models Are Redefining Embodied AI: A Comprehensive Review
Subtle Storm
Subtle Storm
Jan 7, 2026 · Artificial Intelligence

Which Factors Will Define AI in 2026? A Deep Dive into Emerging Trends

The article analyzes how AI in 2026 will shift from conversational hype to actionable agents, featuring paradigm changes toward act‑oriented agents, a split between edge‑efficient and slow‑thinking models, deep multimodal fusion, and embodied intelligence that turns AI into a practical digital colleague.

2026 AI trendsAI agentsedge AI
0 likes · 7 min read
Which Factors Will Define AI in 2026? A Deep Dive into Emerging Trends
HyperAI Super Neural
HyperAI Super Neural
Jan 7, 2026 · Artificial Intelligence

How NASA Engineers and Tech Titans Are Building a $2B General Robot Brain

FieldAI, a 2023 startup backed by Bezos, Gates, Nvidia and Intel, has raised over $405 million to develop a physics‑first “general robot brain” (FFMs) that closes the real‑world data gap, leverages NASA‑honed autonomy research, and targets industrial tasks while riding a surge in global robotics investment.

General-Purpose RobotsNASAembodied AI
0 likes · 11 min read
How NASA Engineers and Tech Titans Are Building a $2B General Robot Brain
21CTO
21CTO
Dec 22, 2025 · Artificial Intelligence

Open-Source XR-1: China’s First Embodied VLA Model for Robots

Beijing Humanoid Robot Innovation Center has open‑sourced XR‑1, the nation’s first VLA (vision‑language‑action) model that meets embodied‑intelligence standards, along with its supporting data sets RoboMIND 2.0 and ArtVIP, detailing its three‑stage training paradigm and cross‑modal capabilities.

ArtVIPOpen SourceRoboMIND
0 likes · 5 min read
Open-Source XR-1: China’s First Embodied VLA Model for Robots
Xiaomi Tech
Xiaomi Tech
Dec 1, 2025 · Artificial Intelligence

Seven Xiaomi AI Papers Accepted at AAAI 2026: Multimodal, Embodied & Database Advances

AAAI 2026 accepted seven Xiaomi research papers—two oral presentations—covering multimodal sound editing, embodied 3D agent scheduling, scalable Text-to-SQL schema linking, parallel speculative decoding, long‑form speech QA, high‑level spatial navigation, and VLM‑driven autonomous‑driving adversaries, each with concrete datasets, methods, and benchmark gains.

AAAI 2026Speech QAText-to-SQL
0 likes · 13 min read
Seven Xiaomi AI Papers Accepted at AAAI 2026: Multimodal, Embodied & Database Advances
Data Party THU
Data Party THU
Nov 16, 2025 · Artificial Intelligence

How X‑VLA Enables 120‑Minute Unassisted Robot Clothing Folding with a 0.9B Model

The X‑VLA paper introduces a 0.9‑billion‑parameter, fully open‑source embodied model that uses a learnable soft‑prompt and divide‑and‑conquer encoding to handle heterogeneous robot vision inputs, achieving a record‑breaking 120‑minute autonomous clothing‑folding task while surpassing benchmarks across five simulation environments.

X-VLAembodied AIflow-matching
0 likes · 7 min read
How X‑VLA Enables 120‑Minute Unassisted Robot Clothing Folding with a 0.9B Model
Amap Tech
Amap Tech
Oct 7, 2025 · Artificial Intelligence

Farsighted-LAM & SSM-VLA: Boosting Spatial‑Temporal Reasoning for Embodied AI

Introducing Farsighted-LAM, a novel latent action model that integrates geometric perception and multi‑scale temporal modeling, and its end‑to‑end SSM‑VLA framework with a Chain‑of‑Thought reasoning module, the authors demonstrate markedly improved spatial‑temporal fidelity, interpretability, and state‑of‑the‑art performance on challenging VLA benchmarks.

chain-of-thoughtembodied AIlatent action models
0 likes · 11 min read
Farsighted-LAM & SSM-VLA: Boosting Spatial‑Temporal Reasoning for Embodied AI
Huawei Cloud Developer Alliance
Huawei Cloud Developer Alliance
Jun 24, 2025 · Artificial Intelligence

Embodied AI Revolution: Key Takeaways from HDC 2025 Roundtable

At Huawei's 2025 Developer Conference in Dongguan, over 120 experts from academia and industry gathered for a roundtable on embodied AI, discussing challenges and breakthroughs in robotics, 3D scene generation, cloud‑edge collaboration, and the future of physical intelligence across sectors.

3D scene generationCloud Computingedge-cloud collaboration
0 likes · 13 min read
Embodied AI Revolution: Key Takeaways from HDC 2025 Roundtable
JD Tech
JD Tech
Jun 20, 2025 · Artificial Intelligence

How JD‑Tech’s AnchorDP3 Dominated the CVPR 2025 Dual‑Arm Robotics Challenge

JD‑Tech leveraged large‑model innovations and a novel AnchorDP3 3D diffusion policy to win both stages of the CVPR 2025 dual‑arm manipulation competition, showcasing breakthroughs in synthetic data generation, multimodal perception, and precise trajectory control for embodied AI robots.

3D diffusion policyCVPR 2025dual-arm manipulation
0 likes · 8 min read
How JD‑Tech’s AnchorDP3 Dominated the CVPR 2025 Dual‑Arm Robotics Challenge
AntTech
AntTech
May 30, 2025 · Artificial Intelligence

Insights from Ant Group’s 10th Technical Open Day: Multimodal, Embodied, and Future Model Architectures for AGI

The Ant Group’s 10th Technical Open Day gathered leading AI experts who examined the current state and future directions of multimodal large models, embodied AI, world models, transformer architectures, and vertical applications, offering a comprehensive view of the challenges and opportunities on the path toward AGI.

AGIAI safetyLarge Language Models
0 likes · 16 min read
Insights from Ant Group’s 10th Technical Open Day: Multimodal, Embodied, and Future Model Architectures for AGI
AntTech
AntTech
Mar 26, 2025 · Artificial Intelligence

BodyGen: A Bio‑Inspired Embodied Co‑Design Framework for Autonomous Robot Evolution

BodyGen, a new embodied co‑design framework presented at ICLR 2025, enables robots to autonomously evolve their morphology and control policies using reinforcement learning and transformer‑based networks, achieving up to 60 % performance gains with a lightweight 1.43 M‑parameter model, and its code is publicly released.

Transformerco-designembodied AI
0 likes · 10 min read
BodyGen: A Bio‑Inspired Embodied Co‑Design Framework for Autonomous Robot Evolution
AI Cyberspace
AI Cyberspace
Feb 23, 2025 · Artificial Intelligence

How Helix Empowers Humanoid Robots to See, Hear, Understand, and Act

Helix is a groundbreaking Vision‑Language‑Action model that integrates perception, language understanding, and motor control, enabling humanoid robots to perform full upper‑body continuous movements, collaborate across multiple robots, grasp any household object via natural language, and run on low‑power embedded GPUs for commercial use.

Humanoid RoboticsVision-Language-Actionembodied AI
0 likes · 16 min read
How Helix Empowers Humanoid Robots to See, Hear, Understand, and Act
Java Tech Enthusiast
Java Tech Enthusiast
Jan 12, 2025 · Artificial Intelligence

AgiBot World: Large-Scale Multi‑Robot Embodied AI Dataset Release

AgiBot World, the first globally‑scale robot dataset captured in fully realistic environments, provides ten‑fold longer trajectories and hundred‑fold greater scene coverage than prior collections, featuring over 80 daily‑life skills recorded by a 32‑DOF robot with advanced sensing, and includes rigorous multi‑stage quality control with future releases slated to reach a million runs and millions of simulated trajectories.

Large Datasetcomputer visionembodied AI
0 likes · 9 min read
AgiBot World: Large-Scale Multi‑Robot Embodied AI Dataset Release
Meituan Technology Team
Meituan Technology Team
Jan 9, 2025 · Artificial Intelligence

Roundtable Discussion on Embodied Intelligence at Meituan Robot Research Institute 2024 Academic Annual Meeting

At Meituan Robot Research Institute’s 2024 academic meeting, a diverse panel of scholars and entrepreneurs debated the relative importance of hardware and algorithms for embodied intelligence, identified near‑term market niches such as hazardous‑environment and household assistance, projected rapid scaling to thousands of autonomous humanoids, and highlighted safety, mass‑market adoption, and ethical considerations as key challenges.

Artificial IntelligenceIndustry Applicationsembodied AI
0 likes · 27 min read
Roundtable Discussion on Embodied Intelligence at Meituan Robot Research Institute 2024 Academic Annual Meeting
AntTech
AntTech
Oct 29, 2024 · Artificial Intelligence

Embodied Intelligence and General‑Purpose Humanoid Robots: Insights from Wang He’s Ant T‑Space Talk

In a detailed presentation, Peking University assistant professor Wang He explained the current state and future direction of embodied intelligence, emphasizing synthetic data, three core intelligences, and the commercial‑grade capabilities of his startup’s general‑purpose humanoid robots across manufacturing, retail, and home applications.

Humanoid Robotembodied AIindustrial automation
0 likes · 17 min read
Embodied Intelligence and General‑Purpose Humanoid Robots: Insights from Wang He’s Ant T‑Space Talk
Architect
Architect
Nov 8, 2023 · Artificial Intelligence

AI Agents Unleashed: From Assistants API to Multi‑Agent Frameworks

The article dissects the rise of AI agents—from OpenAI's Assistants API and multimodal perception‑brain‑action pipelines to retrieval‑augmented generation, tool‑use strategies, single‑ and multi‑agent deployments, and emerging frameworks like AutoGen—while highlighting concrete examples, benchmark results, and current limitations.

AI agentsAssistants APILarge Language Models
0 likes · 38 min read
AI Agents Unleashed: From Assistants API to Multi‑Agent Frameworks
DataFunSummit
DataFunSummit
Nov 4, 2023 · Artificial Intelligence

AIGC Generation Models and Diffusion‑Based Planning for Embodied AI

This article explores powerful AIGC generation models and large language models like ChatGPT, detailing how diffusion models can be applied to robotic planning, introducing AdaptDiffuser, self‑evolving data generation, and embodied AI challenges, while summarizing recent research and practical implementations.

AIGCAdaptDiffuserdiffusion models
0 likes · 20 min read
AIGC Generation Models and Diffusion‑Based Planning for Embodied AI
DataFunTalk
DataFunTalk
Mar 19, 2023 · Artificial Intelligence

Key Technical Directions Highlighted in the GPT‑4 Report and Emerging LLM Research Trends

Zhang Junlin’s answer summarizes the GPT‑4 technical report’s three main research directions—closed‑loop LLM development, capability prediction using small models, and an open LLM evaluation framework—while also noting additional trends such as low‑cost ChatGPT replication and embodied multimodal intelligence.

Capability PredictionGPT-4embodied AI
0 likes · 7 min read
Key Technical Directions Highlighted in the GPT‑4 Report and Emerging LLM Research Trends