Tagged articles

Vision-Language-Action

34 articles · Page 1 of 1
Machine Heart
Machine Heart
Sep 6, 2026 · Artificial Intelligence

VLAct: 16 GPUs, 20% Data Beats GR00T N1.6 in Cross-Embodiment Transfer

VLAct introduces a representation-centric continued pre-training framework for Vision-Language-Action models, achieving 92.5% on RoboTwin 2.0 and surpassing all World Action Models on RoboDojo using only 16 GPUs and open data; with 20% downstream data it outperforms GR00T N1.6 on unseen GR-1 robot.

RoboDojoRoboTwinVLA
0 likes · 10 min read
VLAct: 16 GPUs, 20% Data Beats GR00T N1.6 in Cross-Embodiment Transfer
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 2, 2026 · Artificial Intelligence

How UniSteer Boosted VLA Success from 20% to 90% in Just 66 Minutes

UniSteer introduces a noise‑inversion interface that lets human corrections directly train a lightweight noise actor, enabling a Vision‑Language‑Action model to improve real‑world task success from 20% to 90% within 66 minutes and outperforming DSRL and DAgger baselines.

Human-Guided RLNoise InversionUniSteer
0 likes · 15 min read
How UniSteer Boosted VLA Success from 20% to 90% in Just 66 Minutes
Machine Heart
Machine Heart
Sep 2, 2026 · Artificial Intelligence

How UniSteer Boosts Real‑World VLA Success from 20% to 90% in 66 Minutes

UniSteer introduces a noise‑steering interface that lets human corrections and reinforcement learning jointly update a lightweight noise actor, enabling a Vision‑Language‑Action robot to raise task success from 20% to 90% within 66 minutes while using only two full human demonstrations.

Noise SteeringUniSteerVision-Language-Action
0 likes · 14 min read
How UniSteer Boosts Real‑World VLA Success from 20% to 90% in 66 Minutes
Machine Heart
Machine Heart
Aug 12, 2026 · Artificial Intelligence

A Future‑Predicting Critic Propels VLA Reinforcement Learning

The World Critic Model (WCM) augments the critic in vision‑language‑action reinforcement learning with future state prediction, enabling robots to evaluate not only the current value but also anticipate upcoming dynamics, which dramatically improves both in‑distribution and out‑of‑distribution performance across multiple benchmarks.

OpenMOSSPOMDPVision-Language-Action
0 likes · 12 min read
A Future‑Predicting Critic Propels VLA Reinforcement Learning
Machine Heart
Machine Heart
Aug 10, 2026 · Artificial Intelligence

DriveTeach-VLA Bridges Autonomous Driving Scenes and Foundation Model Pre‑training via Image Trajectories

The ECCV‑2026 paper introduces DriveTeach‑VLA, a vision‑language‑action model that improves autonomous driving by distilling traffic‑aware visual cues and projecting BEV trajectories onto image pixels, achieving state‑of‑the‑art PDMS scores of 90.4 on NAVSIM and up to 92.7 with a trajectory selector, while detailing the training pipeline, visual distillation, 2D‑TGP prompting, and extensive ablations.

Autonomous DrivingVision-Language-Actionfoundation models
0 likes · 11 min read
DriveTeach-VLA Bridges Autonomous Driving Scenes and Foundation Model Pre‑training via Image Trajectories
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 8, 2026 · Artificial Intelligence

How Repositioning the Language Path Boosts VLA Instruction Generalization by 20‑40%

The paper analyzes why Vision‑Language‑Action models fail when task instructions are paraphrased, demonstrates that language semantics remain partially encoded, and shows that the Grounded Semantic Re‑Binding (GSR) redesign of the language‑to‑action information flow improves instruction generalization by up to 40% across multiple VLA architectures.

GSRMultimodal LearningVLA
0 likes · 17 min read
How Repositioning the Language Path Boosts VLA Instruction Generalization by 20‑40%
Machine Heart
Machine Heart
Aug 6, 2026 · Artificial Intelligence

Why the “L” in VLA Isn’t Redundant: Position‑Aware Rebinding Boosts Instruction Generalization 20‑40%

The paper analyzes why Vision‑Language‑Action models fail when instructions are paraphrased, designs a grounded semantic re‑binding (GSR) method that restructures language flow, and demonstrates 20‑40% gains in instruction generalization across several VLA architectures on the LIBERO‑Para benchmark.

GSRLIBERO BenchmarkVLA-Adapter
0 likes · 17 min read
Why the “L” in VLA Isn’t Redundant: Position‑Aware Rebinding Boosts Instruction Generalization 20‑40%
Machine Heart
Machine Heart
Aug 3, 2026 · Artificial Intelligence

Does VLA Action Prediction Need an LLM? TurboVLA Achieves 32 Hz with 0.2 B Params on RTX 4090

TurboVLA, a real‑time vision‑language‑action model from Huazhong University of Science and Technology and Huawei, bypasses the large language model bottleneck by directly fusing visual and language features, achieving 32 Hz online action prediction on a single RTX 4090 with only 0.2 B parameters and 0.9 GB VRAM, while maintaining high success rates across LIBERO, RoboTwin 2.0, and real‑robot tasks.

LIBEROLLMRTX 4090
0 likes · 11 min read
Does VLA Action Prediction Need an LLM? TurboVLA Achieves 32 Hz with 0.2 B Params on RTX 4090
Machine Heart
Machine Heart
Aug 2, 2026 · Artificial Intelligence

Can Tactile Sensing Complete Embodied AI's Perception of the Physical World?

The article examines why vision‑language‑action models alone cannot fully understand physical environments, outlines the limitations of visual‑only embodied AI, and surveys recent algorithmic efforts—such as VTLA, N0‑VTLA, OmniVTLA—to embed tactile perception as a “touch‑neuron” for finer physical interaction.

Multimodal LearningVision-Language-Actionembodied AI
0 likes · 8 min read
Can Tactile Sensing Complete Embodied AI's Perception of the Physical World?
Data Party THU
Data Party THU
Jul 26, 2026 · Artificial Intelligence

Understanding VLA Safety: A Visual Overview and Design Guidelines for Robot Security

The article reviews the Vision‑Language‑Action (VLA) safety landscape, classifies attacks and defenses across training and inference phases, highlights the multimodal attack surface, real‑time constraints, and simulation‑to‑reality gaps, and proposes a fast‑slow dual‑loop defense architecture for safe embodied AI.

AI safetyMultimodal AttackSimulation-to-Reality
0 likes · 9 min read
Understanding VLA Safety: A Visual Overview and Design Guidelines for Robot Security
Machine Heart
Machine Heart
Jul 24, 2026 · Artificial Intelligence

Jetson-PI Enables Real‑Time VLA on Robots, Boosting Jetson Orin Control Frequency 8.66×

The paper presents Jetson-PI, an open‑source VLA real‑time control framework for low‑power edge devices that tackles inference latency and perception‑action misalignment through foresight‑aligned asynchronous correction, confidence‑based scheduling, and edge‑engine optimizations, raising Jetson Orin control frequency from 0.7 Hz to 6.06 Hz and improving task success rates on LIBERO benchmarks and a real‑world clothing‑folding robot.

Asynchronous InferenceJetson OrinReal-time Control
0 likes · 15 min read
Jetson-PI Enables Real‑Time VLA on Robots, Boosting Jetson Orin Control Frequency 8.66×
Machine Heart
Machine Heart
Jul 22, 2026 · Artificial Intelligence

Harness VLA Redefines Embodied AI Execution, Surpassing NVIDIA Cap‑X Performance

Harness VLA introduces a Harness Layer that orchestrates frozen Vision‑Language‑Action models with an Agentic Planner, dramatically improving generalization on challenging robot benchmarks—achieving 82.4% success on LIBERO‑Pro versus 18.2% for NVIDIA Cap‑X—while remaining model‑agnostic and open‑source.

Agentic PlannerVision-Language-Actionbenchmark
0 likes · 21 min read
Harness VLA Redefines Embodied AI Execution, Surpassing NVIDIA Cap‑X Performance
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 9, 2026 · Artificial Intelligence

How Attending Before Acting Boosts Generalization in Pelican-VLA 0.5

The talk presents Pelican-VLA 0.5, a unified Vision‑Language‑Action model that leverages attention‑level generalization without task‑specific supervision, achieving over 91% success on RoboTwin benchmarks and demonstrating early zero‑shot generalization through a novel Reasoning Slots bottleneck.

Attention GeneralizationPelican-VLAVision-Language-Action
0 likes · 6 min read
How Attending Before Acting Boosts Generalization in Pelican-VLA 0.5
Machine Heart
Machine Heart
Jul 2, 2026 · Artificial Intelligence

Quantifying Robot Data Value: ATHENA Scales Influence Functions to Billion‑Parameter VLA with 313× Speedup

ATHENA introduces a data‑curation framework for billion‑parameter multi‑task Vision‑Language‑Action models that extends influence functions via Kronecker gradient compression and a multitask influence interaction scheme, achieving a 313× reduction in compute (from 8054.6 to 25.7 GPU‑hours) and improving task success rates while using fewer, higher‑value demonstrations.

Vision-Language-Actiondata curationinfluence functions
0 likes · 9 min read
Quantifying Robot Data Value: ATHENA Scales Influence Functions to Billion‑Parameter VLA with 313× Speedup
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 30, 2026 · Artificial Intelligence

LabVLA: From Thinking to Doing—What AI Still Needs to Master Scientific Labs

LabVLA introduces a Vision‑Language‑Action paradigm and a knowledge‑enhanced simulation engine to teach AI systems how to plan and execute real‑world scientific experiments, achieving 71.1%/70.0% success in simulated benchmarks and demonstrating comparable performance on a real Franka robot while highlighting remaining challenges for fully autonomous lab assistants.

AI for scienceLabVLARoboGenesis
0 likes · 13 min read
LabVLA: From Thinking to Doing—What AI Still Needs to Master Scientific Labs
Machine Heart
Machine Heart
Jun 28, 2026 · Artificial Intelligence

OpenHLM Enables Whole‑Body Loco‑Manipulation for Humanoid Robots

OpenHLM presents an open‑source VLA recipe that lets humanoid robots coordinate arms, torso, legs, and feet under vision‑language commands, using decoupled whole‑body teleoperation, multi‑step flow generation, and low‑cost data sources to achieve superior whole‑body loco‑manipulation performance on the HLM‑12 benchmark.

Humanoid RoboticsOpenHLMVision-Language-Action
0 likes · 10 min read
OpenHLM Enables Whole‑Body Loco‑Manipulation for Humanoid Robots
Machine Heart
Machine Heart
Jun 26, 2026 · Artificial Intelligence

LabVLA: Bridging AI Reasoning and Hands‑On Lab Automation

LabVLA introduces a vision‑language‑action framework and a knowledge‑enhanced simulation engine to enable AI models to learn and generalize scientific lab manipulation, achieving 71% success on benchmark tasks and demonstrating real‑world performance on a Franka robot, while outlining current limitations and future directions.

AI for scienceLabVLAVision-Language-Action
0 likes · 12 min read
LabVLA: Bridging AI Reasoning and Hands‑On Lab Automation
Machine Heart
Machine Heart
Jun 23, 2026 · Artificial Intelligence

Can VLA‑JEPA Achieve Robust Vision‑Language‑Action with Few Robot Trajectories and Lots of Human Video?

The article analyzes VLA‑JEPA, a JEPA‑style pre‑training framework that combines limited robot trajectories with abundant human video to build a latent world model for Vision‑Language‑Action tasks, showing improved robustness and high success rates across simulated and real‑robot benchmarks.

Self-supervised LearningVLA-JEPAVision-Language-Action
0 likes · 12 min read
Can VLA‑JEPA Achieve Robust Vision‑Language‑Action with Few Robot Trajectories and Lots of Human Video?
Machine Heart
Machine Heart
Jun 10, 2026 · Artificial Intelligence

MINT: Enabling Strong Generalization and One‑Shot Transfer for Vision‑Language‑Action Models

MINT introduces a spectrally disentangled tokenization and intent‑driven strategy that lets Vision‑Language‑Action models generalize compositionally, transfer with a single demonstration, and achieve state‑of‑the‑art performance and robustness across benchmark suites and real‑world robot experiments.

Few-shot TransferMINTVision-Language-Action
0 likes · 9 min read
MINT: Enabling Strong Generalization and One‑Shot Transfer for Vision‑Language‑Action Models
Machine Heart
Machine Heart
Jun 5, 2026 · Artificial Intelligence

Beyond Binary Success: Redefining Fine-Grained Manipulation Evaluation for Embodied AI

The paper introduces MetaFine, a diagnostic meta‑evaluation framework that moves robot manipulation assessment from a simple success/failure binary to a three‑dimensional analysis of understanding, perception, and behavior, revealing up to 70% over‑estimation in traditional benchmarks and offering a hybrid real‑sim testing pipeline for fair, reproducible results.

MetaFineVision-Language-Actiondiagnostic evaluation
0 likes · 12 min read
Beyond Binary Success: Redefining Fine-Grained Manipulation Evaluation for Embodied AI
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Jun 2, 2026 · Artificial Intelligence

Halving Training Time: LoongForge Full‑Stack Optimizations Boost GR00T N1.6 Throughput 2.3×

LoongForge applies system‑level optimizations—async data prefetch, fine‑grained communication‑compute overlap via a Megatron distributed optimizer, and per‑microbatch CUDA Graph scheduling—to the GR00T N1.6 Vision‑Language‑Action model, delivering up to 2.3× higher training throughput and a 56.6% reduction in overall training time on an 8×A800 cluster.

CUDA GraphDistributed TrainingGR00T N1.6
0 likes · 14 min read
Halving Training Time: LoongForge Full‑Stack Optimizations Boost GR00T N1.6 Throughput 2.3×
Machine Heart
Machine Heart
May 28, 2026 · Artificial Intelligence

How AutoMoT Leverages Large‑Model Understanding for End‑to‑End Driving Decisions and Trajectory Planning

AutoMoT introduces a unified Vision‑Language‑Action model that combines a 4B Qwen3‑VL understanding expert with a 1.6B action expert via layer‑wise shared attention and asynchronous inference, achieving state‑of‑the‑art results on Bench2Drive and nuScenes while preserving general VLM capabilities.

Asynchronous InferenceAutoMoTAutonomous Driving
0 likes · 10 min read
How AutoMoT Leverages Large‑Model Understanding for End‑to‑End Driving Decisions and Trajectory Planning
Machine Heart
Machine Heart
May 22, 2026 · Artificial Intelligence

Can World Action Models Replace VLA? Nvidia’s New Embodied AI Paradigm Reviewed

The article reviews the emerging World Action Model (WAM) paradigm, critiques the limitations of Vision‑Language‑Action models, outlines cascaded and joint WAM architectures, discusses required data sources, evaluation metrics, and future challenges, positioning WAM as a new foundational approach for embodied AI.

Future State PredictionVision-Language-Actiondata fusion
0 likes · 14 min read
Can World Action Models Replace VLA? Nvidia’s New Embodied AI Paradigm Reviewed
Machine Heart
Machine Heart
May 22, 2026 · Artificial Intelligence

HiF-VLA: Motion‑Centric ‘Think‑While‑Doing’ World Action Model Breaks Short‑Sighted Limits

HiF-VLA introduces a motion‑centric bidirectional spatiotemporal reasoning framework with a joint‑expert module that simultaneously predicts future visual motion and generates high‑precision action sequences, eliminating visual redundancy, cutting inference latency and memory usage, and achieving superior success rates on long‑horizon benchmarks such as CALVIN and LIBERO‑LONG.

HiF-VLAMotion RepresentationVision-Language-Action
0 likes · 9 min read
HiF-VLA: Motion‑Centric ‘Think‑While‑Doing’ World Action Model Breaks Short‑Sighted Limits
Machine Heart
Machine Heart
May 16, 2026 · Artificial Intelligence

Why Robots Need World Models: A Joint Survey from Leading Institutions

This article surveys recent advances in robot world models, explaining why predictive models are essential for embodied intelligence, how they integrate with Vision‑Language‑Action systems, the various architectural approaches, benchmark trends, and the remaining challenges for reliable deployment.

Vision-Language-Actionbenchmarkrobot learning
0 likes · 14 min read
Why Robots Need World Models: A Joint Survey from Leading Institutions
Meituan Technology Team
Meituan Technology Team
Apr 23, 2026 · Artificial Intelligence

LARYBench Introduces an ImageNet‑Style Benchmark for Embodied Action Representations Learned from Human Video

LARYBench (Latent Action Representation Yielding Benchmark) provides the first systematic, ImageNet‑scale evaluation for implicit action representations derived from large‑scale human video, decoupling representation quality from downstream control, and shows that general‑purpose vision models outperform specialized embodied models in both action generalization and control precision across diverse robot morphologies and environments.

Vision-Language-Actionaction representationbenchmark
0 likes · 13 min read
LARYBench Introduces an ImageNet‑Style Benchmark for Embodied Action Representations Learned from Human Video
Machine Heart
Machine Heart
Apr 18, 2026 · Artificial Intelligence

Eliminating ‘Think‑Then‑Act’ Stalls: StreamingVLA Boosts VLA Speed by 2.4×

StreamingVLA introduces action‑flow matching and adaptive early observation to parallelize generation, execution, and perception in vision‑language‑action models, cutting per‑action latency from 49.9 ms to 31.6 ms, reducing stall time 6.5‑fold, and achieving up to 2.4× end‑to‑end speedup in LIBERO benchmarks and real‑world robot tests.

LIBEROStreamingVLAVision-Language-Action
0 likes · 13 min read
Eliminating ‘Think‑Then‑Act’ Stalls: StreamingVLA Boosts VLA Speed by 2.4×
Machine Heart
Machine Heart
Apr 11, 2026 · Artificial Intelligence

Why VLA Pioneers Are Abandoning Vision‑Language‑Action Models

Generalist AI’s GEN-1 model achieves over 99% success, 2‑3× speed gains with only a tenth of the data, and its founders argue that vision‑language‑action (VLA) models are merely a crutch, urging a shift toward goal‑driven, fully‑scratch training for physical AGI.

GEN-1Generalist AIGoal-driven research
0 likes · 13 min read
Why VLA Pioneers Are Abandoning Vision‑Language‑Action Models
Machine Heart
Machine Heart
Mar 31, 2026 · Artificial Intelligence

Point‑VLA: Overcoming Embodied AI’s Language Bottleneck with Visual Grounding

The Point‑VLA method introduced by Qianxun AI’s Gaoyang team tackles the fundamental limits of language‑only instruction in vision‑language‑action models by adding visual grounding via bounding‑box cues, boosting real‑robot success rates from 32.4% to 92.5% across six challenging tasks.

Multimodal LearningPoint-VLAVision-Language-Action
0 likes · 13 min read
Point‑VLA: Overcoming Embodied AI’s Language Bottleneck with Visual Grounding
HyperAI Super Neural
HyperAI Super Neural
Feb 19, 2026 · Artificial Intelligence

World Model & VLA Breakthroughs: Top Papers from NVIDIA, ByteDance, Tsinghua and Others

This roundup highlights six recent embodied AI papers that advance world models and vision‑language‑action (VLA) techniques, covering DreamDojo's massive first‑person video model, LingBot‑World simulator, Agent World Model generator, BagelVLA, ACoT‑VLA, and the closed‑loop World‑VLA‑Loop framework.

Synthetic EnvironmentsVision-Language-Actionembodied AI
0 likes · 8 min read
World Model & VLA Breakthroughs: Top Papers from NVIDIA, ByteDance, Tsinghua and Others
HyperAI Super Neural
HyperAI Super Neural
Dec 12, 2025 · Artificial Intelligence

Weekly AI Paper Digest: Attention, Nvidia VLA, TTS, and Graph Neural Networks

This roundup presents five recent AI papers covering hierarchical sparse attention for ultra‑long context, Nvidia's Alpamayo‑R1 VLA model for autonomous driving, the non‑autoregressive F5‑TTS system, LatentMAS for latent‑space multi‑agent collaboration, and Deeper‑GXX that deepens arbitrary graph neural networks, highlighting each method's key innovations and reported performance gains.

Attention MechanismAutonomous DrivingGraph Neural Networks
0 likes · 6 min read
Weekly AI Paper Digest: Attention, Nvidia VLA, TTS, and Graph Neural Networks
Data Party THU
Data Party THU
Oct 29, 2025 · Artificial Intelligence

Can Test-Time Scaling Unlock More Reliable Vision‑Language‑Action Robots?

The paper introduces RoboMonkey, a framework that applies a generate‑and‑verify paradigm and test‑time scaling to Vision‑Language‑Action models, showing that increasing sampling and verification at inference dramatically reduces action error across multiple VLA architectures, and presents scalable verifier training, synthetic data augmentation, and efficient deployment strategies.

AI researchAction VerificationRoboMonkey
0 likes · 8 min read
Can Test-Time Scaling Unlock More Reliable Vision‑Language‑Action Robots?
Amap Tech
Amap Tech
Oct 6, 2025 · Artificial Intelligence

Breaking VLA Training Limits: World-Env’s Virtual Sandbox for Safe, Data‑Efficient Robotics

World-Env introduces a virtual training sandbox that eliminates physical interaction, dramatically improves data efficiency with just five expert demos per task, and employs a vision‑language model as a semantic judge to dynamically terminate actions, enabling safe, high‑performing VLA post‑training across diverse robotic benchmarks.

Data EfficiencyVision-Language-Actionvirtual environment
0 likes · 9 min read
Breaking VLA Training Limits: World-Env’s Virtual Sandbox for Safe, Data‑Efficient Robotics
AI Cyberspace
AI Cyberspace
Feb 23, 2025 · Artificial Intelligence

How Helix Empowers Humanoid Robots to See, Hear, Understand, and Act

Helix is a groundbreaking Vision‑Language‑Action model that integrates perception, language understanding, and motor control, enabling humanoid robots to perform full upper‑body continuous movements, collaborate across multiple robots, grasp any household object via natural language, and run on low‑power embedded GPUs for commercial use.

Humanoid RoboticsVision-Language-Actionembodied AI
0 likes · 16 min read
How Helix Empowers Humanoid Robots to See, Hear, Understand, and Act