ABot-M0.5: The First Unified World Action Model for Mobile Manipulation
ABot-M0.5 introduces a unified world action model that aligns video prediction, intermediate latent actions, and low‑level physical control for mobile manipulation, achieving state‑of‑the‑art long‑horizon success rates and fine‑grained precision across benchmarks such as RoboCasa365, RoboTwin, and LIBERO, while detailing novel architectural components and a three‑stage progressive training regime.
