LeaP: Learning Where Robot Action Generation Should Begin

LeaP introduces a learnable source prior conditioned on proprioception that jointly predicts mean and variance of a diagonal Gaussian for initializing generative robot policies, achieving 81.6% average success on 15 RoboTwin tasks (25.5 pp over standard Gaussian) and 80% on real Franka tasks (33.3 pp gain), with faster convergence.

Machine Heart
Machine Heart
Machine Heart
LeaP: Learning Where Robot Action Generation Should Begin

Generative robot policies such as diffusion models and flow matching models treat action generation as conditional sampling from expert demonstrations. These policies typically start from a standard Gaussian distribution independent of the current observation. Recent methods like A2A and VITA use proprioceptive or visual information to construct a deterministic starting point, but do not explicitly model its uncertainty.

Researchers from Southeast University propose LeaP (Learnable source Prior) , a proprioception-driven learnable source prior that jointly predicts the mean and variance of a diagonal Gaussian distribution. This provides a state-dependent stochastic initialization for action generation. The prior head is a two-layer MLP that takes proprioceptive features and outputs mean and log-variance. During sampling, standard Gaussian noise is scaled per dimension and shifted by the predicted mean. Both training and inference retain this stochastic sampling; each timestep shares distribution parameters but samples independently, leaving temporal structure to the downstream generator.

The generator receives both visual and proprioceptive features, while the prior receives only proprioception. This separation lets LeaP plug into existing generators without changing their architecture or inference solver. Training uses three joint objectives: a flow matching loss to learn the transformation from source samples to expert actions, a negative log-likelihood loss to supervise the conditional probability distribution of the prior, and a contrastive alignment loss to enforce feature-space correspondence between source samples and paired expert actions. The prior and generator are optimized end-to-end.

Experiments

Experiments run on 15 dual-arm manipulation tasks in RoboTwin (pick-and-place, tool use, bimanual coordination) with 50 expert demonstrations per task, 100 evaluations per seed, and 3 seeds. All methods share a DP3 PointNet point-cloud encoder. The direct baseline NoPrior uses the same encoder, generator, and flow matching pipeline but with a standard Gaussian start. LeaP achieves 81.6% average success rate , a 25.5 percentage-point improvement over NoPrior. It also outperforms deterministic-start methods VITA (+6.5 pp) and A2A (+7.8 pp).

Model size: LeaP totals 13.22M parameters, with the prior head contributing only 0.21M (~1.6%). On the “open laptop” task, LeaP reaches 91% success after 1,000 epochs, surpassing four major baselines trained for 3,000 epochs, indicating better convergence efficiency.

Real-robot validation on a Franka Research 3 arm (pick cube, close box, pick-and-place sandbag) with 100 teleoperated demos and 20 randomized test trials per task yields 80.0% average success for LeaP, versus 68.3% (A2A), 56.7% (VITA), and 46.7% (NoPrior) — a 33.3 pp gain over NoPrior .

Ablation Studies

Three ablation dimensions are examined on three RoboTwin tasks (pick different bottles, open laptop, handover block), each evaluated 100 times.

Prior input: Proprioception-only LeaP achieves 85.3% success. Adding visual features or fusing vision with proprioception drops performance to 59.3%–70.3%. Visualization shows proprioceptive priors place samples near target actions, while vision-injected variants drift away. This supports the division of labor: prior sets an action-relevant start from robot state; generator refines with vision.

Distribution form: All learnable priors beat the standard Gaussian baseline. However, full-covariance Gaussian and Gaussian mixture models do not exceed LeaP’s diagonal Gaussian. A three-way comparison: predicted mean only → 78.0%; mean + fixed unit-variance Gaussian noise → 62.7%; joint learning of mean and state-adaptive variance → 85.3%. This shows that mere stochasticity is insufficient; learning state-adaptive uncertainty is more effective than adding fixed-magnitude noise.

Training supervision: Training the prior with only flow matching loss gives 59.7% (vs. 47.7% without prior). Adding likelihood supervision raises it to 74.3%; adding contrastive alignment raises it to 74.7%; full model reaches 85.3%. Likelihood loss models conditional density of expert actions; contrastive loss constrains feature-space correspondence — they provide complementary signals.

Cross-generator generality: The same prior lifts flow matching from 47.7% to 85.3%. In a diffusion bridge framework, replacing a deterministic start (prior mean) with the full LeaP distribution improves success from 68.7% to 76.7%, confirming that learning state-adaptive variance adds value even when a state-dependent mean exists.

Limitations noted: evaluation focuses on tabletop manipulation; diagonal Gaussian does not explicitly model inter-dimensional action correlations; more complex prior forms and robot morphologies remain for future work. The paper is accepted at CoRL 2026; code is open-sourced at https://github.com/SEU-VIPGroup/LeaP and project page at https://daimeipo.github.io/LeaP/.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Flow MatchingImitation LearningRobot ManipulationLeaPCoRL 2026Generative Robot PoliciesProprioceptive Priors
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.