Tagged articles

ECCV 2026

16 articles · Page 1 of 1
vivo Internet Technology
vivo Internet Technology
Sep 23, 2026 · Artificial Intelligence

ART: Two-Stage Makeup Transfer Anchors Supervision to Real Images, Beats Pseudo-Target Ceiling

vivo BlueImage Lab introduces ART, a two-stage makeup transfer framework that first learns from synthetic pseudo-targets then refines using real reference images as supervision, achieving state-of-the-art fidelity on complex makeup like glitter and face painting while preserving identity, and releases the first 2K makeup dataset MF2K.

ECCV 2026MF2K datasetcomputer vision
0 likes · 15 min read
ART: Two-Stage Makeup Transfer Anchors Supervision to Real Images, Beats Pseudo-Target Ceiling
Machine Heart
Machine Heart
Sep 20, 2026 · Artificial Intelligence

HSImul3R: Physics-in-the-Loop Reconstruction Turns Human Videos into Robot Skills

HSImul3R introduces a physics-in-the-loop framework that reconstructs simulation-ready human-scene interactions from sparse views, using scene-targeted reinforcement learning and direct simulation reward optimization to achieve stable physical interactions, validated on HSIBench and deployed on Unitree G1 robot.

3D ReconstructionECCV 2026HSIBench
0 likes · 11 min read
HSImul3R: Physics-in-the-Loop Reconstruction Turns Human Videos into Robot Skills
Machine Heart
Machine Heart
Sep 15, 2026 · Artificial Intelligence

REAL: Embodied Agents Navigate Open Worlds Without Oracle Perception or Perfect Instructions

Researchers from Shanghai Jiao Tong University and Shanghai AI Lab introduce REAL, an ECCV 2026 framework that enables embodied agents to actively explore unknown environments, disambiguate vague user instructions through dialogue, and execute mobile manipulation tasks via a unified MCP tool interface, achieving 78.3% real-world success on a dual-arm robot.

Active ExplorationECCV 2026Embodied AI
0 likes · 19 min read
REAL: Embodied Agents Navigate Open Worlds Without Oracle Perception or Perfect Instructions
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 12, 2026 · Artificial Intelligence

How Robots Turn World Representations into Action: Jiajun Wu's ECCV 2026 Insights

Stanford professor Jiajun Wu's ECCV 2026 talk explores how structured world representations enable robots to act in complex physical environments, covering compositional skill learning, neural kinematics for zero-shot object interaction, video diffusion models for generating free demonstrations, and new benchmarks for evaluating embodied reasoning.

BenchmarksECCV 2026Embodied AI
0 likes · 35 min read
How Robots Turn World Representations into Action: Jiajun Wu's ECCV 2026 Insights
Machine Heart
Machine Heart
Sep 10, 2026 · Artificial Intelligence

AgentVLN: VLM-as-Brain Architecture for Agentic Robot Navigation at ECCV 2026

AgentVLN introduces a VLM-as-Brain architecture for vision-language navigation, using cross-space representation mapping to translate 3D paths into 2D visual prompts, context-driven self-correction for error recovery, and query-driven perceptual chain-of-thought for active information gathering, achieving state-of-the-art results on R2R-CE and RxR-CE benchmarks with a 3B model deployable on Jetson edge devices.

AgentVLNCross-Space Representation MappingECCV 2026
0 likes · 10 min read
AgentVLN: VLM-as-Brain Architecture for Agentic Robot Navigation at ECCV 2026
AntTech
AntTech
Sep 2, 2026 · Artificial Intelligence

ECCV 2026 Paper Showcase: Breaking Vision Perception & Prediction Boundaries

This article summarizes four ECCV 2026 papers: a dynamic cross-layer injection framework for deep vision-language fusion, an event-augmented VLA model enabling robot operation in extreme darkness and blur, a predictive differentiable rendering method using 2D Gaussians for high-fidelity video prediction, and a data-scaling approach for high-resolution weather forecasting that demonstrates clear scaling laws.

Data ScalingECCV 2026Embodied AI
0 likes · 10 min read
ECCV 2026 Paper Showcase: Breaking Vision Perception & Prediction Boundaries
Data Party THU
Data Party THU
Sep 1, 2026 · Artificial Intelligence

How OVOW Turns Monocular Video into Physically Simulatable 4D Meshes

OVOW introduces a pipeline that converts ordinary monocular video into instance‑level 4D meshes with real‑world scale, motion, and physical properties, enabling direct editing and simulation in physics engines while outperforming prior 4D reconstruction methods in accuracy and speed.

4D reconstructionECCV 2026OVOW
0 likes · 9 min read
How OVOW Turns Monocular Video into Physically Simulatable 4D Meshes
Machine Heart
Machine Heart
Aug 19, 2026 · Artificial Intelligence

Geometry‑Grounded Video Diffusion for 3D‑Consistent World Generation

DreamWorld introduces a two‑stage Geometry‑then‑Appearance pipeline that first generates geometry features for the target viewpoint using a 3D foundation model and then conditions a video diffusion model on these features to produce high‑quality, 3D‑consistent RGB videos, achieving state‑of‑the‑art results on multiple benchmarks.

3D GeometryECCV 2026Geometry-Appearance Decoupling
0 likes · 12 min read
Geometry‑Grounded Video Diffusion for 3D‑Consistent World Generation
Machine Heart
Machine Heart
Aug 5, 2026 · Artificial Intelligence

Atomic Dance: Explainable Music-to-Dance Generation Using Atomic Movements

The paper introduces Atomic Dance, a two‑stage framework that first discovers repeatable, semantically labeled atomic movements and then plans and completes dance sequences, achieving more coherent, rhythm‑aligned and editable music‑driven choreography, as demonstrated on the AIST++ benchmark.

Atomic MovementsChoreographyECCV 2026
0 likes · 9 min read
Atomic Dance: Explainable Music-to-Dance Generation Using Atomic Movements
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 20, 2026 · Artificial Intelligence

How VGGRPO Uses 4D Latent Rewards for World‑Consistent Video Generation

VGGRPO introduces a latent‑space geometry model and two 4D rewards—camera motion smoothness and geometry reprojection consistency—to improve geometric consistency in video diffusion models without sacrificing pre‑training generalization, achieving smoother camera paths and coherent scene structures even in dynamic scenarios.

4D rewardECCV 2026Reinforcement Learning
0 likes · 9 min read
How VGGRPO Uses 4D Latent Rewards for World‑Consistent Video Generation
Machine Heart
Machine Heart
Jul 17, 2026 · Artificial Intelligence

VGGRPO: 4D Latent Rewards for World‑Consistent Video Generation (ECCV 2026)

VGGRPO introduces a latent‑space geometry model and two 4D rewards—camera‑motion smoothness and geometry‑reprojection consistency—to eliminate drift and improve structural coherence in video diffusion models without altering their pretrained architecture, achieving state‑of‑the‑art results on static and dynamic benchmarks.

4D rewardECCV 2026Reinforcement Learning
0 likes · 7 min read
VGGRPO: 4D Latent Rewards for World‑Consistent Video Generation (ECCV 2026)
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jul 15, 2026 · Artificial Intelligence

Is Video Generation the ‘Next Token Prediction’ for Vision? Insights from the GenCeption Paper

The ECCV 2026 paper by He Kaiming, Andrew Zisserman and collaborators proposes GenCeption, a text‑to‑video model repurposed as a single‑forward DiT‑based vision learner that unifies depth, segmentation, pose and 3D tasks, and evaluates its multi‑task performance against specialized baselines.

DiTECCV 2026GenCeption
0 likes · 10 min read
Is Video Generation the ‘Next Token Prediction’ for Vision? Insights from the GenCeption Paper
Machine Heart
Machine Heart
Jul 11, 2026 · Artificial Intelligence

Real-Time Multi-Shot Long Video Generation: Introducing ShotStream (ECCV 2026)

ShotStream tackles the high latency and zero‑interaction problems of multi‑shot long video generation by proposing a streaming architecture with a dual‑cache memory, discontinuous RoPE, and a two‑stage self‑forcing distillation, achieving over 25× speedup to 16 FPS on a single H200 GPU and outperforming existing bidirectional and autoregressive models.

ECCV 2026ShotStreamdual cache
0 likes · 8 min read
Real-Time Multi-Shot Long Video Generation: Introducing ShotStream (ECCV 2026)
Machine Heart
Machine Heart
Jul 7, 2026 · Artificial Intelligence

Unlocking Free‑View Video Virtual Try‑On with TryOnCrafter’s 4D Try‑On Proxy

TryOnCrafter introduces a camera‑controllable video virtual try‑on framework that builds a renderable 4D proxy to enable free‑view, 360° and bullet‑time effects while preserving structural stability, texture consistency, and realistic motion across arbitrary camera trajectories.

3D Gaussian Splatting4D try-on proxyECCV 2026
0 likes · 16 min read
Unlocking Free‑View Video Virtual Try‑On with TryOnCrafter’s 4D Try‑On Proxy
Machine Heart
Machine Heart
Jul 4, 2026 · Artificial Intelligence

LinStereo Bridges the Last Mile of Stereo Matching (ECCV 2026)

LinStereo replaces ConvGRU with a position‑aware linear attention module, adds a multi‑scale cost volume and monocular depth initialization, cutting Middlebury occlusion error by 37%, outperforming larger models, and achieving strong zero‑shot underwater performance while remaining parameter‑efficient.

ECCV 2026Linear AttentionStereo Matching
0 likes · 10 min read
LinStereo Bridges the Last Mile of Stereo Matching (ECCV 2026)
Amap Tech
Amap Tech
Jun 30, 2026 · Artificial Intelligence

Six ECCV 2026 Papers – Vision, Video Generation, Visual‑Language Navigation

ECCV 2026 received 10,473 submissions and accepted 2,883 (27.5%); Gaode contributed six papers spanning computer vision, generative video, and visual‑language navigation, each presenting novel reinforcement‑learning or multimodal frameworks, new datasets, and benchmark results that outperform prior state‑of‑the‑art methods.

ECCV 2026Reinforcement Learningcomputer vision
0 likes · 13 min read
Six ECCV 2026 Papers – Vision, Video Generation, Visual‑Language Navigation