How 30,000 Hours of Tactile Data Give Embodied AI a Real “Sense of Touch”
NeoteAI and Fudan University release a 30,000‑hour visual‑tactile dataset and three models—NeoForce, VTLA and TWAM—that demonstrate tactile scaling, predictive touch for VLA, and multimodal world‑model integration, achieving up to 99% task success and proving touch as a core building block for embodied intelligence.
NeoteAI, in collaboration with Fudan University, announced three technical reports that elevate tactile perception from an auxiliary modality to core infrastructure for embodied AI. The centerpiece is the NeoData dataset, which aggregates more than 30,000 hours of synchronized visual‑tactile interaction video, roughly 1.4 million operation clips (33 billion timesteps), 80 billion RGB frames and 100 billion tactile frames collected by six robot platforms (Franka, Piper, UR5e, etc.) across 450 real‑world tasks performed by 90 operators. Five thousand hours of the video have been open‑sourced.
The NeoForce unified representation model learns sensor‑agnostic, temporally structured tactile embeddings that can be transferred to downstream embodied‑intelligence policies, effectively creating a “tactile Mandarin” that bridges heterogeneous hardware.
Building on the dataset, the VTLA (Vision‑Language‑Action) approach replaces the conventional lagging tactile feedback with a predictive module that forecasts tactile signals for the next 50 steps. By encoding current tactile input into compact tokens and conditioning on visual context, VTLA anticipates slip and force trends, enabling pre‑emptive motion adjustments. Empirical results show an 85% success rate on plug‑insertion (vs. 60% for vision‑only) and 99% on key‑removal (vs. 35% vision‑only). Moreover, VTLA leverages failure data via a progress‑assessment model and offline reinforcement learning, raising success rates on towel‑folding, backpack packing, and box‑folding from 50%/35%/20% to 95%/80%/75% respectively.
The TWAM (Tactile‑World‑Action‑Model) extends world‑model generation by jointly predicting future video, tactile, and action sequences. It employs a mixed‑expert architecture with three asynchronous expert networks (video, tactile, action). The video expert is warm‑started from a pretrained model, while tactile and action experts use narrower structures, resulting in a total of 7.2 billion parameters—about half of a full‑width baseline. In simulation, TWAM achieves an average success rate of 84.5% (baseline 36%); on real robots, it attains 46.3% across eight tasks (LingBot‑VA baseline 21.9%). Ablation studies confirm that removing either future tactile prediction or current tactile conditioning degrades performance markedly, underscoring tactile’s indispensable role in world‑model reasoning.
Collectively, the three reports demonstrate that tactile modality exhibits scaling potential comparable to vision, can be seamlessly integrated into existing VLA pipelines, and should be treated as a first‑class modality in world‑model design. While challenges remain—data‑collection cost, cross‑platform generalization, and real‑time inference—the released resources (dataset, models, code) constitute a significant step toward robots that not only see but also feel the world.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
