TouchScale: 500-Hour Unified Vision-Tactile Dataset Shows Scaling Benefits Transfer to Robot Control

Researchers from 15 institutions introduce TouchScale, a 500-hour human vision-tactile dataset collected with a single wearable system, demonstrating that scaling tactile data under consistent hardware and procedures yields cross-sensor generalization gains and lifts real-robot manipulation success from 22.5% to 57.5% across four contact-rich tasks without using human action labels.

Machine Heart
Machine Heart
Machine Heart
TouchScale: 500-Hour Unified Vision-Tactile Dataset Shows Scaling Benefits Transfer to Robot Control

Dataset Construction: 500 Hours of Unified Vision-Tactile Data

TouchScale comprises 500 hours of synchronized first-person vision and touch data collected by a single wearable system across ~20 participants, ~87,000 interaction recordings, ~2,000 task descriptions, and >1,500 objects in 9 environment types and 800+ scene configurations. The hardware suite includes a head-mounted RGB-D camera for scene context, two wrist-mounted RGB cameras for close-up hand-object interaction, and two full-hand tactile gloves each containing 880 taxels (tactile sensing elements) covering five fingers and palm to capture normal pressure during contact.

Quality Control: Hardware and Algorithmic Filtering

To ensure high-quality tactile signals at scale, TouchScale employs low-noise flexible skin sensors at acquisition time. Post-collection, a VLM-assisted automatic quality-check pipeline cross-references wrist-camera video (identifying grasp, release, and contacting fingers) with actual tactile responses. Clear sensor failures are discarded automatically; ambiguous cases (single-finger dropout, severe occlusion, insufficient visual evidence) are routed for human review.

Cross-Sensor Tactile Generalization at Fixed Data Scale

In zero-shot cross-sensor tactile prediction, models trained on TouchScale subsets (≈16 hours) achieve contact cIoU of 0.181 when evaluated on the EgoTactile dataset (different glove layout, results mapped to 12 anatomical hand regions), outperforming EgoTouch-trained models (cIoU 0.134) at the same training volume. This isolates the benefit of unified collection from mere data quantity.

Tactile Supervision Improves Visual Representations

Using tactile signals as supervision for visual encoder pretraining (same initialization and compute budget), TouchScale-pretrained encoders surpass those pretrained on OpenTouch, FEEL, and EgoTouch on linear-probe and end-to-end fine-tuning across three benchmarks: MECCANO, Something-Something V2, and Ego-Exo4D action recognition.

Real-Robot Manipulation: Mid-Training on Human Vision-Tactile Data

Experiments use an xArm6 arm with a BrainCo Revo 2 dexterous hand on four contact-intensive tasks: soft/hard object sorting, bottle-cap removal, test-tube transfer, and whiteboard wiping. Each task receives 50 robot demonstrations and 20 test rollouts. Both conditions start from the same N0-VTLA checkpoint, use identical vision-tactile inputs, and share the same post-training on robot demos. The only difference: one condition adds TouchScale mid-training (future tactile prediction from current vision, language, and touch) before robot post-training. No human action labels or hand-to-robot motion mapping are used.

Results show average success rising from 22.5% to 57.5%, with every task improving:

Soft/hard sorting: 10% → 60%

Bottle-cap removal: 40% → 70%

Test-tube transfer: 30% → 60%

Whiteboard wiping: 10% → 40%

Scaling Analysis: Benefits Continue Beyond Perception to Control

In tactile prediction, holding model and evaluation fixed while scaling TouchScale from 10% to 100% raises cIoU from 0.311 to 0.383, confirming continued cross-sensor generalization gains. The same scaling trend transfers to robot control: with robot demonstration data held constant, increasing TouchScale mid-training data drives the average success rate from 22.5% to 57.5%.

Limitations and Outlook

Current real-robot validation covers only one hardware platform and four contact-rich tasks. The mechanism by which human tactile experience translates to improved robot control remains an open question. Nevertheless, TouchScale demonstrates that tactile data can follow the same scaling trajectory as first-person video: systematic, large-scale collection under fixed sensing and procedure enables learning transferable physical interaction priors from unlabeled human experience.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Embodied AIrobot manipulationmid-trainingcross-sensor generalizationdataset scalingtactile predictionTouchScalevision-tactile learning
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.