Tagged articles

Vision-Language-Action

25 articles · Page 1 of 1
Data Party THU
Data Party THU
Jul 26, 2026 · Artificial Intelligence

Understanding VLA Safety: A Visual Overview and Design Guidelines for Robot Security

The article reviews the Vision‑Language‑Action (VLA) safety landscape, classifies attacks and defenses across training and inference phases, highlights the multimodal attack surface, real‑time constraints, and simulation‑to‑reality gaps, and proposes a fast‑slow dual‑loop defense architecture for safe embodied AI.

AI safetyDefense StrategiesMultimodal Attack
0 likes · 9 min read
Understanding VLA Safety: A Visual Overview and Design Guidelines for Robot Security
Machine Heart
Machine Heart
Jul 24, 2026 · Artificial Intelligence

Jetson-PI Enables Real‑Time VLA on Robots, Boosting Jetson Orin Control Frequency 8.66×

The paper presents Jetson-PI, an open‑source VLA real‑time control framework for low‑power edge devices that tackles inference latency and perception‑action misalignment through foresight‑aligned asynchronous correction, confidence‑based scheduling, and edge‑engine optimizations, raising Jetson Orin control frequency from 0.7 Hz to 6.06 Hz and improving task success rates on LIBERO benchmarks and a real‑world clothing‑folding robot.

Asynchronous InferenceJetson OrinReal-time Control
0 likes · 15 min read
Jetson-PI Enables Real‑Time VLA on Robots, Boosting Jetson Orin Control Frequency 8.66×
Machine Heart
Machine Heart
Jul 22, 2026 · Artificial Intelligence

Harness VLA Redefines Embodied AI Execution, Surpassing NVIDIA Cap‑X Performance

Harness VLA introduces a Harness Layer that orchestrates frozen Vision‑Language‑Action models with an Agentic Planner, dramatically improving generalization on challenging robot benchmarks—achieving 82.4% success on LIBERO‑Pro versus 18.2% for NVIDIA Cap‑X—while remaining model‑agnostic and open‑source.

Agentic PlannerGeneralizationVision-Language-Action
0 likes · 21 min read
Harness VLA Redefines Embodied AI Execution, Surpassing NVIDIA Cap‑X Performance
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 9, 2026 · Artificial Intelligence

How Attending Before Acting Boosts Generalization in Pelican-VLA 0.5

The talk presents Pelican-VLA 0.5, a unified Vision‑Language‑Action model that leverages attention‑level generalization without task‑specific supervision, achieving over 91% success on RoboTwin benchmarks and demonstrating early zero‑shot generalization through a novel Reasoning Slots bottleneck.

Attention GeneralizationPelican-VLAVision-Language-Action
0 likes · 6 min read
How Attending Before Acting Boosts Generalization in Pelican-VLA 0.5
Machine Heart
Machine Heart
Jul 2, 2026 · Artificial Intelligence

Quantifying Robot Data Value: ATHENA Scales Influence Functions to Billion‑Parameter VLA with 313× Speedup

ATHENA introduces a data‑curation framework for billion‑parameter multi‑task Vision‑Language‑Action models that extends influence functions via Kronecker gradient compression and a multitask influence interaction scheme, achieving a 313× reduction in compute (from 8054.6 to 25.7 GPU‑hours) and improving task success rates while using fewer, higher‑value demonstrations.

Vision-Language-Actiondata curationinfluence functions
0 likes · 9 min read
Quantifying Robot Data Value: ATHENA Scales Influence Functions to Billion‑Parameter VLA with 313× Speedup
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 30, 2026 · Artificial Intelligence

LabVLA: From Thinking to Doing—What AI Still Needs to Master Scientific Labs

LabVLA introduces a Vision‑Language‑Action paradigm and a knowledge‑enhanced simulation engine to teach AI systems how to plan and execute real‑world scientific experiments, achieving 71.1%/70.0% success in simulated benchmarks and demonstrating comparable performance on a real Franka robot while highlighting remaining challenges for fully autonomous lab assistants.

AI for ScienceLabVLARoboGenesis
0 likes · 13 min read
LabVLA: From Thinking to Doing—What AI Still Needs to Master Scientific Labs
Machine Heart
Machine Heart
Jun 28, 2026 · Artificial Intelligence

OpenHLM Enables Whole‑Body Loco‑Manipulation for Humanoid Robots

OpenHLM presents an open‑source VLA recipe that lets humanoid robots coordinate arms, torso, legs, and feet under vision‑language commands, using decoupled whole‑body teleoperation, multi‑step flow generation, and low‑cost data sources to achieve superior whole‑body loco‑manipulation performance on the HLM‑12 benchmark.

Humanoid RoboticsOpenHLMVision-Language-Action
0 likes · 10 min read
OpenHLM Enables Whole‑Body Loco‑Manipulation for Humanoid Robots
Machine Heart
Machine Heart
Jun 26, 2026 · Artificial Intelligence

LabVLA: Bridging AI Reasoning and Hands‑On Lab Automation

LabVLA introduces a vision‑language‑action framework and a knowledge‑enhanced simulation engine to enable AI models to learn and generalize scientific lab manipulation, achieving 71% success on benchmark tasks and demonstrating real‑world performance on a Franka robot, while outlining current limitations and future directions.

AI for ScienceLabVLAVision-Language-Action
0 likes · 12 min read
LabVLA: Bridging AI Reasoning and Hands‑On Lab Automation
Machine Heart
Machine Heart
Jun 23, 2026 · Artificial Intelligence

Can VLA‑JEPA Achieve Robust Vision‑Language‑Action with Few Robot Trajectories and Lots of Human Video?

The article analyzes VLA‑JEPA, a JEPA‑style pre‑training framework that combines limited robot trajectories with abundant human video to build a latent world model for Vision‑Language‑Action tasks, showing improved robustness and high success rates across simulated and real‑robot benchmarks.

VLA-JEPAVision-Language-Actionbenchmark
0 likes · 12 min read
Can VLA‑JEPA Achieve Robust Vision‑Language‑Action with Few Robot Trajectories and Lots of Human Video?
Machine Heart
Machine Heart
Jun 10, 2026 · Artificial Intelligence

MINT: Enabling Strong Generalization and One‑Shot Transfer for Vision‑Language‑Action Models

MINT introduces a spectrally disentangled tokenization and intent‑driven strategy that lets Vision‑Language‑Action models generalize compositionally, transfer with a single demonstration, and achieve state‑of‑the‑art performance and robustness across benchmark suites and real‑world robot experiments.

Compositional GeneralizationFew-shot TransferMINT
0 likes · 9 min read
MINT: Enabling Strong Generalization and One‑Shot Transfer for Vision‑Language‑Action Models
Machine Heart
Machine Heart
Jun 5, 2026 · Artificial Intelligence

Beyond Binary Success: Redefining Fine-Grained Manipulation Evaluation for Embodied AI

The paper introduces MetaFine, a diagnostic meta‑evaluation framework that moves robot manipulation assessment from a simple success/failure binary to a three‑dimensional analysis of understanding, perception, and behavior, revealing up to 70% over‑estimation in traditional benchmarks and offering a hybrid real‑sim testing pipeline for fair, reproducible results.

MetaFineVision-Language-Actiondiagnostic evaluation
0 likes · 12 min read
Beyond Binary Success: Redefining Fine-Grained Manipulation Evaluation for Embodied AI
Baidu Intelligent Cloud Tech Hub
Baidu Intelligent Cloud Tech Hub
Jun 2, 2026 · Artificial Intelligence

Halving Training Time: LoongForge Full‑Stack Optimizations Boost GR00T N1.6 Throughput 2.3×

LoongForge applies system‑level optimizations—async data prefetch, fine‑grained communication‑compute overlap via a Megatron distributed optimizer, and per‑microbatch CUDA Graph scheduling—to the GR00T N1.6 Vision‑Language‑Action model, delivering up to 2.3× higher training throughput and a 56.6% reduction in overall training time on an 8×A800 cluster.

CUDA GraphDistributed TrainingGR00T N1.6
0 likes · 14 min read
Halving Training Time: LoongForge Full‑Stack Optimizations Boost GR00T N1.6 Throughput 2.3×
Machine Heart
Machine Heart
May 28, 2026 · Artificial Intelligence

How AutoMoT Leverages Large‑Model Understanding for End‑to‑End Driving Decisions and Trajectory Planning

AutoMoT introduces a unified Vision‑Language‑Action model that combines a 4B Qwen3‑VL understanding expert with a 1.6B action expert via layer‑wise shared attention and asynchronous inference, achieving state‑of‑the‑art results on Bench2Drive and nuScenes while preserving general VLM capabilities.

Asynchronous InferenceAutoMoTBench2Drive
0 likes · 10 min read
How AutoMoT Leverages Large‑Model Understanding for End‑to‑End Driving Decisions and Trajectory Planning
Machine Heart
Machine Heart
May 22, 2026 · Artificial Intelligence

Can World Action Models Replace VLA? Nvidia’s New Embodied AI Paradigm Reviewed

The article reviews the emerging World Action Model (WAM) paradigm, critiques the limitations of Vision‑Language‑Action models, outlines cascaded and joint WAM architectures, discusses required data sources, evaluation metrics, and future challenges, positioning WAM as a new foundational approach for embodied AI.

Future State PredictionVision-Language-Actiondata fusion
0 likes · 14 min read
Can World Action Models Replace VLA? Nvidia’s New Embodied AI Paradigm Reviewed
Machine Heart
Machine Heart
May 22, 2026 · Artificial Intelligence

HiF-VLA: Motion‑Centric ‘Think‑While‑Doing’ World Action Model Breaks Short‑Sighted Limits

HiF-VLA introduces a motion‑centric bidirectional spatiotemporal reasoning framework with a joint‑expert module that simultaneously predicts future visual motion and generates high‑precision action sequences, eliminating visual redundancy, cutting inference latency and memory usage, and achieving superior success rates on long‑horizon benchmarks such as CALVIN and LIBERO‑LONG.

HiF-VLAMotion RepresentationVision-Language-Action
0 likes · 9 min read
HiF-VLA: Motion‑Centric ‘Think‑While‑Doing’ World Action Model Breaks Short‑Sighted Limits
Machine Heart
Machine Heart
May 16, 2026 · Artificial Intelligence

Why Robots Need World Models: A Joint Survey from Leading Institutions

This article surveys recent advances in robot world models, explaining why predictive models are essential for embodied intelligence, how they integrate with Vision‑Language‑Action systems, the various architectural approaches, benchmark trends, and the remaining challenges for reliable deployment.

Vision-Language-ActionWorld Modelsbenchmark
0 likes · 14 min read
Why Robots Need World Models: A Joint Survey from Leading Institutions
Meituan Technology Team
Meituan Technology Team
Apr 23, 2026 · Artificial Intelligence

LARYBench Introduces an ImageNet‑Style Benchmark for Embodied Action Representations Learned from Human Video

LARYBench (Latent Action Representation Yielding Benchmark) provides the first systematic, ImageNet‑scale evaluation for implicit action representations derived from large‑scale human video, decoupling representation quality from downstream control, and shows that general‑purpose vision models outperform specialized embodied models in both action generalization and control precision across diverse robot morphologies and environments.

Vision-Language-Actionaction representationbenchmark
0 likes · 13 min read
LARYBench Introduces an ImageNet‑Style Benchmark for Embodied Action Representations Learned from Human Video
Machine Heart
Machine Heart
Apr 18, 2026 · Artificial Intelligence

Eliminating ‘Think‑Then‑Act’ Stalls: StreamingVLA Boosts VLA Speed by 2.4×

StreamingVLA introduces action‑flow matching and adaptive early observation to parallelize generation, execution, and perception in vision‑language‑action models, cutting per‑action latency from 49.9 ms to 31.6 ms, reducing stall time 6.5‑fold, and achieving up to 2.4× end‑to‑end speedup in LIBERO benchmarks and real‑world robot tests.

LIBEROParallel ExecutionStreamingVLA
0 likes · 13 min read
Eliminating ‘Think‑Then‑Act’ Stalls: StreamingVLA Boosts VLA Speed by 2.4×
Machine Heart
Machine Heart
Apr 11, 2026 · Artificial Intelligence

Why VLA Pioneers Are Abandoning Vision‑Language‑Action Models

Generalist AI’s GEN-1 model achieves over 99% success, 2‑3× speed gains with only a tenth of the data, and its founders argue that vision‑language‑action (VLA) models are merely a crutch, urging a shift toward goal‑driven, fully‑scratch training for physical AGI.

GEN-1Generalist AIGoal-driven research
0 likes · 13 min read
Why VLA Pioneers Are Abandoning Vision‑Language‑Action Models
Machine Heart
Machine Heart
Mar 31, 2026 · Artificial Intelligence

Point‑VLA: Overcoming Embodied AI’s Language Bottleneck with Visual Grounding

The Point‑VLA method introduced by Qianxun AI’s Gaoyang team tackles the fundamental limits of language‑only instruction in vision‑language‑action models by adding visual grounding via bounding‑box cues, boosting real‑robot success rates from 32.4% to 92.5% across six challenging tasks.

Point-VLAVision-Language-ActionVisual Grounding
0 likes · 13 min read
Point‑VLA: Overcoming Embodied AI’s Language Bottleneck with Visual Grounding
HyperAI Super Neural
HyperAI Super Neural
Feb 19, 2026 · Artificial Intelligence

World Model & VLA Breakthroughs: Top Papers from NVIDIA, ByteDance, Tsinghua and Others

This roundup highlights six recent embodied AI papers that advance world models and vision‑language‑action (VLA) techniques, covering DreamDojo's massive first‑person video model, LingBot‑World simulator, Agent World Model generator, BagelVLA, ACoT‑VLA, and the closed‑loop World‑VLA‑Loop framework.

Synthetic EnvironmentsVision-Language-ActionWorld Models
0 likes · 8 min read
World Model & VLA Breakthroughs: Top Papers from NVIDIA, ByteDance, Tsinghua and Others
HyperAI Super Neural
HyperAI Super Neural
Dec 12, 2025 · Artificial Intelligence

Weekly AI Paper Digest: Attention, Nvidia VLA, TTS, and Graph Neural Networks

This roundup presents five recent AI papers covering hierarchical sparse attention for ultra‑long context, Nvidia's Alpamayo‑R1 VLA model for autonomous driving, the non‑autoregressive F5‑TTS system, LatentMAS for latent‑space multi‑agent collaboration, and Deeper‑GXX that deepens arbitrary graph neural networks, highlighting each method's key innovations and reported performance gains.

Multi-Agent SystemsVision-Language-Actionattention mechanism
0 likes · 6 min read
Weekly AI Paper Digest: Attention, Nvidia VLA, TTS, and Graph Neural Networks
Data Party THU
Data Party THU
Oct 29, 2025 · Artificial Intelligence

Can Test-Time Scaling Unlock More Reliable Vision‑Language‑Action Robots?

The paper introduces RoboMonkey, a framework that applies a generate‑and‑verify paradigm and test‑time scaling to Vision‑Language‑Action models, showing that increasing sampling and verification at inference dramatically reduces action error across multiple VLA architectures, and presents scalable verifier training, synthetic data augmentation, and efficient deployment strategies.

AI researchAction VerificationRoboMonkey
0 likes · 8 min read
Can Test-Time Scaling Unlock More Reliable Vision‑Language‑Action Robots?
Amap Tech
Amap Tech
Oct 6, 2025 · Artificial Intelligence

Breaking VLA Training Limits: World-Env’s Virtual Sandbox for Safe, Data‑Efficient Robotics

World-Env introduces a virtual training sandbox that eliminates physical interaction, dramatically improves data efficiency with just five expert demos per task, and employs a vision‑language model as a semantic judge to dynamically terminate actions, enabling safe, high‑performing VLA post‑training across diverse robotic benchmarks.

Data EfficiencyVision-Language-ActionWorld Model
0 likes · 9 min read
Breaking VLA Training Limits: World-Env’s Virtual Sandbox for Safe, Data‑Efficient Robotics
AI Cyberspace
AI Cyberspace
Feb 23, 2025 · Artificial Intelligence

How Helix Empowers Humanoid Robots to See, Hear, Understand, and Act

Helix is a groundbreaking Vision‑Language‑Action model that integrates perception, language understanding, and motor control, enabling humanoid robots to perform full upper‑body continuous movements, collaborate across multiple robots, grasp any household object via natural language, and run on low‑power embedded GPUs for commercial use.

Humanoid RoboticsVision-Language-Actionembodied AI
0 likes · 16 min read
How Helix Empowers Humanoid Robots to See, Hear, Understand, and Act