Tagged articles

large-scale dataset

8 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 17, 2026 · Artificial Intelligence

How RynnWorld‑4D Gives Robots a 3D‑Future Vision with 4D World Modeling

The article analyzes the limitations of 2D video‑based world models for robotic manipulation, introduces the RGB‑DF 4D representation and a three‑branch Transformer with joint cross‑modal attention, details a staged training pipeline and a massive 4D dataset, and demonstrates superior geometry, motion, and policy performance on real‑world dual‑arm tasks.

4D world modelRGB-DF representationdiffusion model
0 likes · 14 min read
How RynnWorld‑4D Gives Robots a 3D‑Future Vision with 4D World Modeling
Machine Heart
Machine Heart
Jul 8, 2026 · Artificial Intelligence

SurgMotion: Billion‑Parameter Model Pushes Video AI to Motion Prediction

SurgMotion, the world’s first billion‑parameter surgical video foundation model trained on the 15‑million‑frame, 3,658‑hour SurgMotion‑15M dataset, introduces motion‑guided latent masking, spatiotemporal affinity self‑distillation and feature‑diversity regularization, delivering up to 16.5% gains in workflow recognition and 2.9% error reduction on static tasks while topping 17 benchmark evaluations and becoming the most‑downloaded surgical model on Hugging Face.

SurgMotionfoundation modellarge-scale dataset
0 likes · 8 min read
SurgMotion: Billion‑Parameter Model Pushes Video AI to Motion Prediction
Machine Heart
Machine Heart
Jul 6, 2026 · Artificial Intelligence

How a Million‑Scale 159‑Category Dataset and Foundation Model Set New Standards for Remote‑Sensing Object Detection

The paper introduces LEVIRDet‑159, the largest unified remote‑sensing detection dataset with 159 categories and 2.5 M annotations, and its foundation model LEVIRDetNet, which achieves state‑of‑the‑art performance on nine external benchmarks after a single training run, demonstrating strong cross‑scene generalization.

LEVIRDetObject Detectioncross-benchmark evaluation
0 likes · 9 min read
How a Million‑Scale 159‑Category Dataset and Foundation Model Set New Standards for Remote‑Sensing Object Detection
JD Retail Technology
JD Retail Technology
Jun 29, 2026 · Artificial Intelligence

Uni-AdGen: Unified Autoregressive Model for Personalized Image‑Text Ad Generation (CVPR 2026)

Uni‑AdGen unifies image and text generation in a single autoregressive framework, introduces a coarse‑to‑fine preference module and foreground‑aware control, and demonstrates superior performance on the million‑scale PAd1M dataset with novel evaluation metrics for personalized advertising.

Autoregressive Modelevaluation metricslarge-scale dataset
0 likes · 15 min read
Uni-AdGen: Unified Autoregressive Model for Personalized Image‑Text Ad Generation (CVPR 2026)
Machine Heart
Machine Heart
Apr 11, 2026 · Artificial Intelligence

How 100,000 Hours of Human Data Propelled Psi‑R2 to Lead MolmoSpaces

Lingchu AI demonstrates that scaling human‑operation data to nearly 100,000 hours, combined with a two‑model system and reinforcement learning, can replace costly robot‑teleoperation data and achieve top performance on the MolmoSpaces benchmark.

Embodied AIPsi-R2Psi-W0
0 likes · 12 min read
How 100,000 Hours of Human Data Propelled Psi‑R2 to Lead MolmoSpaces
Amap Tech
Amap Tech
Apr 14, 2025 · Artificial Intelligence

HumanRig: Learning Automatic Rigging for Humanoid Characters Using a Large‑Scale Dataset

HumanRig introduces a large‑scale dataset of 11,434 AI‑generated T‑pose humanoid meshes with unified skeletons, skinning weights, joint data and images, and leverages it in a novel automatic rigging pipeline—featuring a prior‑guided skeleton estimator, a U‑shaped point transformer, and a mesh‑skeleton mutual attention network—that significantly outperforms previous methods in skeleton accuracy and skinning quality.

3D animationAIautomatic rigging
0 likes · 12 min read
HumanRig: Learning Automatic Rigging for Humanoid Characters Using a Large‑Scale Dataset
NewBeeNLP
NewBeeNLP
May 29, 2024 · Artificial Intelligence

How Ant’s Multimodal Team Boosted Video‑Text Retrieval by 24% and Cut Copyright Search Costs 85%

This article presents Ant Group's multimodal research on video retrieval, detailing a large Chinese video‑text pre‑training dataset, three techniques that raise video‑text semantic search performance by up to 24.5%, and an end‑to‑end video‑video copyright detection system that reduces storage by 85% and speeds up inference 18‑fold.

copyright detectionfine-grained modelinghard sample mining
0 likes · 40 min read
How Ant’s Multimodal Team Boosted Video‑Text Retrieval by 24% and Cut Copyright Search Costs 85%
AIWalker
AIWalker
Jan 22, 2024 · Artificial Intelligence

Depth Anything: An Open-Source Large-Scale Model for Arbitrary Image Depth Estimation

Depth Anything introduces a highly practical monocular depth estimation model that leverages a 62‑million‑image unlabeled dataset, teacher‑student training, strong data perturbations, and DINOv2‑based semantic supervision to achieve zero‑shot capability and state‑of‑the‑art performance over MiDaS across multiple benchmarks.

Computer VisionDINOv2Depth Estimation
0 likes · 8 min read
Depth Anything: An Open-Source Large-Scale Model for Arbitrary Image Depth Estimation