World Synesthesia Model Achieves Robust Dexterous In-Hand Manipulation Under Real-World Perturbations
Sharpa Robotics' World Synesthesia Model (WSM), accepted at CoRL 2026, unifies visual geometry, tactile contact, proprioception, and action history into a reusable world model state, enabling a 22-DoF five-finger hand to achieve robust, generalizable in-hand rotation across unseen objects and real-world perturbations through clean depth supervision and recurrent memory.
The Sharpa Robotics team presents the World Synesthesia Model (WSM), a unified framework for robust and generalizable dexterous in-hand manipulation accepted at CoRL 2026. The core challenge addressed is moving dexterous hands from controlled demos to real-world deployment where noisy depth, sparse tactile sensing, external perturbations, and unseen objects cause existing policies to degrade into near-open-loop finger routines.
1. Four Core Modules Form the Unified World Synesthesia Model
The system centers on a Dreamer-style Recurrent State Space Model (RSSM) combined with clean depth reconstruction, action-conditioned recurrent reasoning, PPO asymmetric actor-critic, and a reusable pretrained prior. This operates on a human-scale 22-DoF five-finger hand (Sharpa Wave platform) achieving perturbation-robust, object-ID-agnostic in-hand rotation.
1.1 Multimodal Predictive State Encoding: Visual Geometry + Tactile Contact + Action History
WSM encodes proprioception, tactile, and wrist depth separately, fuses them, and applies noise and dropout to depth inputs during training while targeting clean wrist depth reconstruction. This forces the recurrent state to recover usable hand-object geometry from noisy observations. At deployment, the actor receives both current observations and the WSM recurrent feature w_t = h_t, continuously inferring object pose, geometry, contact, and slip under partial observability.
1.2 World Synesthesia Model: Clean Depth Supervision + Recurrent Memory Are Both Essential
Ablation experiments reveal the performance gain stems from coupling clean geometry supervision with action-conditioned recurrent reasoning. Key results: noisy depth supervision return 708.0 vs full WSM 753.3 (+45.3); removing recurrent input drops to 705.4; frame-wise latent policy only 667.6. Clean geometry plus recurrent memory is the optimal combination.
1.3 WSM Latent Depth Reconstruction: Recovering Hand-Object Geometry from Noisy Input
Real hardware wrist depth streams suffer from high noise and incompleteness due to depth camera precision degradation at close range and NaN-induced holes. The WSM latent variable reconstructs cleaner hand-object geometry, as shown in side-by-side GIF comparisons of real noisy wrist depth versus WSM predicted depth.
1.4 Reusable Synesthetic Prior: 9-Object Pretraining Transfers to 49 New Objects
A nine-object z-axis pretrained WSM serves as a reusable physics prior. After 3000 epochs, average rotation reaches 9.37 ± 0.13 rad/episode (vs 3.28 rad without prior), and drop rate falls from 6% to 0.3%. t-SNE visualization of WSM recurrent states shows structured clustering, and downstream evaluation on 49 objects demonstrates strong zero-shot transfer.
1.5 Sim-to-Real Closed-Loop Deployment: Real Platform Perturbation-Robust Rotation
On the real Sharpa Wave platform, the policy demonstrates continuous multi-object rollout, out-of-distribution (OOD) recovery, and perturbation recovery. GIFs show the hand maintaining stable rotation over one minute, recovering after objects are pushed toward the palm (3-6 seconds to readjust), and reorienting challenging initial poses (e.g., bulb reoriented with three fingers before rotation).
2. Multi-Object Z-Axis Rotation: One Policy Covers Nine Daily Objects
The policy rotates nine diverse objects (duck, strawberry, corner block >1 minute continuous rotation) without object-specific tuning, demonstrating object-agnostic capability.
3. Multi-Axis Rotation, Zero-Shot Generalization, and Tool Use
3.1 Y-Axis Rotation: Tool-Like Objects Sim-to-Real
Y-axis rotation involves larger moment arms and asymmetric lateral torques. Real hardware validates closed-loop y-axis rotation on elongated objects like toothpaste and screwdriver.
3.2 Gravity-Invariant Rotation: Maintaining Grasp Under Palm Pose Changes
Compact grasp keeps objects securely in hand, maintaining contact and stable rotation even when palm orientation changes.
3.3 Zero-Shot Unseen Objects + Process Perturbations + Challenging Initial Poses
The nine-object z-axis policy directly evaluates three unseen geometries zero-shot. Under process perturbations, a cross block pushed toward the palm readjusts to a comfortable workspace in 3-6 seconds. Challenging initial poses are first corrected (e.g., bulb reoriented with three fingers) then rotated.
3.4 Tool Use: Goal-Conditioned Translation + High-Speed Axial Rotation
Demonstrates goal-conditioned translation and high-speed in-hand rotation, extending beyond pure rotation to functional manipulation.
3.5 Real-World Comparison with Mainstream Baselines
Real z-axis corner block rotation comparison: open-loop replay (no feedback, cannot recover controlled rotation), Touch Dexterity (pure tactile, frequent object tipping), In-Hand Rotation (noisy depth causes OOD instability), versus WM-Craftnet (stable continuous rotation >1 minute). The recurrent task context of WSM enables sustained stable rotation where baselines fail.
4. A New Path Toward Real-World Generalization for Dexterous Manipulation
This CoRL 2026 work achieves cross-object, cross-rotation-axis, perturbation-recoverable general in-hand rotation on a 22-DoF hand. Beyond task completion, it answers a key question: can pretrained world models truly serve real dexterous manipulation and improve policy generalization and sim-to-real transfer?
First, reusable synesthetic prior: the world model acts as a general physics prior for downstream policies, enabling efficient transfer and scale-up. WSM learns transferable multimodal physical representations across vision, touch, proprioception, and action history — not single-object or trajectory-specific encodings. The prior helps policies quickly adapt to 49 new objects, showing generalization need not rely solely on larger end-to-end policies but can come from reusable world representations.
Second, sim-to-real deployment: the world model processes noisy depth in real time, closing the vision gap. Real in-hand manipulation suffers from close-range noise, occlusion, and holes in wrist depth. WSM leverages clean depth supervision and action-conditioned recurrent memory to recover reliable hand-object geometric state from noisy observations, enabling unseen object rotation and perturbation recovery on real hardware.
Third, a crucial know-how: world models need not "imagine the future" — they can first learn to "understand the present" better. WM-Craftnet does not primarily use the world model for long-horizon future prediction but as a multimodal world state representation fusing proprioception, touch, noisy depth, and action history into a more complete, robust policy context. Its gains on unseen objects and sim-to-real suggest that for real robots, the practical value of world models may lie not in dreaming the future but in learning a better representation of the present.
Overall, this work proves that vision, touch, proprioception, and action history can be unified into a reusable, noise-robust, physically meaningful world state representation that further serves policy learning and real deployment. It provides a practical direction for world models in robotics: shifting from pursuing "future prediction" toward "building better present-world representations" to support stable generalization of dexterous manipulation across more objects and real perturbation conditions.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
