LIFT: Teaching VLA Models Force Perception via Post-Training Without Force Pretraining
Shanghai Jiao Tong University's LIFT method enables Vision-Language-Action models to learn force perception through post-training alone, achieving significant performance gains on contact-rich tasks with only 20-30 force-labeled demonstrations while preserving pretrained knowledge via weight copying and shifted causal attention.
Shanghai Jiao Tong University researchers (Lu Cewu and Wen Chuan teams) with Qongche propose LIFT (Late Reactive Injection of Force for VLA Post-Training), a method that teaches Vision-Language-Action (VLA) models force perception solely through post-training, without any force-labeled pretraining data.
Problem: Force Data Scarcity and Modality Integration
Force sensing is critical for contact-rich manipulation (e.g., book insertion, ring placement) where visual cues barely change during contact. However, force data collection is expensive and tightly bound to specific hardware, making large-scale pretraining impractical. Adding a new modality also risks catastrophic forgetting of pretrained spatial reasoning and generalization.
LIFT Approach
Reactive Force Injection for High-Frequency Control
LIFT duplicates the pretrained action expert into a "reactive action expert" that attends to recent force history via causal cross-attention, adjusting actions at 10 Hz while vision-language context is precomputed and cached (up from 1 Hz).
Preserving Pretrained Knowledge via Weight Copying and Shifted Causal Attention
At initialization, the reactive expert's weights are copied from the base expert. A shifted causal attention mask ensures the reactive tokens receive identical context as base tokens, aligning outputs numerically and preserving priors.
Two-Stage Training with Mixed Visual and Force Data
Stage 1: Collect force-free visual data using RoboPocket to learn task basics. Stage 2: Deploy on robot, gather 20–30 force-labeled corrections via human intervention. During stage 2, visual-only samples have their force gradients masked to avoid overfitting.
Experimental Results
Evaluated on three tasks:
Towel folding: 73.3 → 84.2 success rate
Book insertion: 36.7 → 58.3
Hanoi ring placement: 26.7 → 56.7
Only 20–30 online force episodes (2–3 hours) were needed. Force improved even non-contact-heavy tasks (towel folding) and accelerated learning.
Four Key Conclusions
Post-training alone suffices for force learning – no force-pretraining required.
Reactive injection is critical – enables within-chunk corrections; especially vital for dense-contact tasks.
Online data collection matches policy distribution – mitigates force distribution shift caused by minor position changes.
Pretrained knowledge is retained – generalization to cloth, lighting, and object changes remains intact after post-training.
Code and details: https://arxiv.org/abs/2607.14236, https://lift-policy.github.io/, https://github.com/y-wng/lift.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
