Tagged articles

vision-language navigation

5 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 12, 2026 · Artificial Intelligence

Robots Retracing LLMs' Scaling Path: LightNav-0 & Light REACT Explained

Light Source Innovation, founded by ex-OpenAI RLHF expert Jiang Xu, releases LightNav-0 for zero-shot cross-morphology navigation and Light REACT for whole-body resilience control, applying LLM-style scalable pre-training, alignment, and deployment paradigms to embodied AI with sim-to-real synthetic data and preference-aligned RL.

RLHFembodied AIpreference alignment
0 likes · 14 min read
Robots Retracing LLMs' Scaling Path: LightNav-0 & Light REACT Explained
Machine Heart
Machine Heart
Sep 10, 2026 · Artificial Intelligence

AgentVLN: VLM-as-Brain Architecture for Agentic Robot Navigation at ECCV 2026

AgentVLN introduces a VLM-as-Brain architecture for vision-language navigation, using cross-space representation mapping to translate 3D paths into 2D visual prompts, context-driven self-correction for error recovery, and query-driven perceptual chain-of-thought for active information gathering, achieving state-of-the-art results on R2R-CE and RxR-CE benchmarks with a 3B model deployable on Jetson edge devices.

AgentVLNCross-Space Representation MappingECCV 2026
0 likes · 10 min read
AgentVLN: VLM-as-Brain Architecture for Agentic Robot Navigation at ECCV 2026
Machine Heart
Machine Heart
Apr 30, 2026 · Artificial Intelligence

Can Internet Videos Replace 3D Annotations? Introducing SceneVerse++ – the Largest Real‑World 3D Scene Dataset

The BIGAI team presents SceneVerse++, a massive real‑world indoor 3D scene dataset built from unlabelled internet videos via an automated pipeline, and demonstrates substantial zero‑shot and fine‑tuned performance gains on 3D detection, spatial VQA, and vision‑language navigation tasks.

3D scene understandingSceneVerse++automated data pipeline
0 likes · 18 min read
Can Internet Videos Replace 3D Annotations? Introducing SceneVerse++ – the Largest Real‑World 3D Scene Dataset
Amap Tech
Amap Tech
Oct 4, 2025 · Artificial Intelligence

How JanusVLN Redefines Vision‑Language Navigation with Dual Implicit Memory

JanusVLN presents a groundbreaking Vision‑and‑Language Navigation framework that decouples semantic understanding from spatial geometry using dual implicit memory, eliminates explicit memory overhead, achieves state‑of‑the‑art performance with only RGB video input, and dramatically improves efficiency and generalization across VLN benchmarks.

3D spatial reasoningDual Implicit Memorymultimodal LLM
0 likes · 10 min read
How JanusVLN Redefines Vision‑Language Navigation with Dual Implicit Memory
Meituan Technology Team
Meituan Technology Team
Jun 15, 2023 · Artificial Intelligence

Meituan Technical Team's 8 CVPR 2023 Papers: Overview and Insights

This article reviews eight CVPR 2023 papers selected by Meituan’s technology team, covering self‑supervised learning, domain adaptation, federated learning, object detection, 3D reconstruction, GAN‑based pre‑training, RGB‑T tracking, vision‑language navigation, and visual‑textual layout generation, highlighting each work’s methodology, experiments, and reported performance gains.

3D Object DetectionCVPR 2023Computer Vision
0 likes · 15 min read
Meituan Technical Team's 8 CVPR 2023 Papers: Overview and Insights