Big Data 15 min read

Alibaba Cloud's Agentic Lake + Streaming: Multimodal Data Platform for AI Agents

At Yunqi 2026, Alibaba Cloud unveiled its Agentic Lake + Agentic Streaming architecture, combining DLF, Flink, EMR Ray/Daft, StarRocks, and Milvus to provide a unified multimodal data infrastructure for AI agents in autonomous driving, embodied intelligence, and enterprise operations.

Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Alibaba Cloud's Agentic Lake + Streaming: Multimodal Data Platform for AI Agents

Four Transformations Driving Data Platform Evolution

Alibaba Cloud's Wang Feng outlined four shifts reshaping big data platforms in the Agentic AI era: data structure expands from text to full multimodal (images, audio, video, sensor signals, 3D geometry); analytics operators upgrade from BI statistics to AI inference; management moves from closed data warehouses to open lakehouses; and users change from human analysts to AI Agents. This reframes data platforms as intelligent infrastructure serving both humans and Agents.

Three Design Principles

Design for AI scenarios

Open-source ecosystem with open architecture

Unified management of multimodal data

Five core products coordinate to cover the complete processing chain: DLF manages multimodal storage and metadata; Flink handles stream processing and Streaming Agents; EMR Ray/Daft provides offline multimodal AI compute; StarRocks enables real-time hybrid retrieval; Milvus serves as the vector engine.

DLF: Building the Multimodal Data Lake with Lake-Stream Unification

DLF uses Apache Paimon as the core lake format. Its Omni Catalog delivers unified schema, catalog, and permission management across OSS object storage (images, audio, video), RDS, SLS, real-time video streams (RTSP/RTMP/HLS), and document knowledge bases (DingTalk, Yuque, PDF). Storage features include cache acceleration, hot-cold tiering, lifecycle management, adaptive bucketing, and file compaction, complemented by lineage, index, and sort optimization.

Lake-Stream Unification architecture: The lake table layer (Paimon) provides minute-level auto-ingestion; the real-time layer (Fluss) offers millisecond-level changelog write-and-consume. Union Read merges historical snapshots with latest increments seamlessly, letting one dataset serve both offline and real-time workloads.

Paimon evolution: Paimon 1.0 targeted large-scale offline lakehouses for e-commerce, retail, logistics. Paimon + Fluss added second-level real-time for finance and gaming. Paimon 2.0 upgrades to full multimodal management , natively supporting LeRobot, RLDS, HDF5 formats, focusing on autonomous driving and embodied intelligence. Flink acts as the streaming ingestion pipeline to keep lake data fresh.

DLF architecture diagram
DLF architecture diagram

EMR Ray + Daft: Distributed Multimodal AI Compute Pipeline

After data lands in the lake, EMR Ray + Daft handles vectorization, annotation, quality assessment, and format conversion. EMR Ray clusters provide independently elastic CPU and GPU resource pools. Daft DataFrame acts as the multimodal AI operator engine with vectorized execution, locality scheduling, and faulty GPU detection/isolation. Dual interfaces support Python DataFrame and SQL paradigms.

Built-in algorithm libraries target embodied intelligence (scene reconstruction, hand reconstruction, action annotation), autonomous driving (image annotation, vectorization, index creation), and LLM data processing (semantic deduplication, quality scoring, tokenization). Deep integration with Paimon multimodal lake format enables joint optimization of AI and relational operators for Data+AI unification.

Case study: Woji Technology's Ego-Centric hand 3D reconstruction pipeline. From first-person monocular video, the pipeline reconstructs dual-hand 3D motion trajectories across four stages: video preprocessing & calibration, depth & hand reconstruction, long-video merging & trajectory cleaning, data export. It involves 10+ steps including GeoCalib intrinsics estimation, MoGe-2 depth estimation, HaWoR dual-hand reconstruction, DROID-SLAM camera tracking. EMR Ray delivers four engineering advantages: model residency reuse, independent CPU/GPU elasticity, multi-video pipeline execution, and checkpoint restart. Output is LeRobot v3 format dataset with three-panel visualization, ready for embodied intelligence model training.

EMR Ray + Daft pipeline
EMR Ray + Daft pipeline

EMR StarRocks: Real-Time Hybrid Retrieval for Multimodal Data

StarRocks evolves along two axes: extreme structured analytics performance and multimodal hybrid retrieval. It simultaneously provides scalar analysis, vector search, full-text search, and AI inference. Federated query supports StarRocks native tables and DLF lake tables (Paimon, Lance, Iceberg), covering real-time BI to AI RAG and AI Memory.

Qwen multimodal data loop: Alibaba's ALake stores multimodal data in Paimon wide tables — each row links ID, attributes, text, image, audio, video, and vectors. StarRocks powers three stages: (1) Collection: hybrid retrieval + Agent multi-round iteration for Top-K sampling evaluation, producing relevance and coverage reports. (2) Training preparation: qualified datasets exported via Top-M large-scale retrieval as specialized training sets for Qwen multimodal training, achieving end-to-end closure.

Autonomous driving case: Zhurui Technology's smart driving data preparation. Their multimodal lake holds video clips, driving frames & target crops, LiDAR point clouds, vehicle logs, geo info, historical faults, and visual vectors. DLF manages metadata; Lance lake tables stored on OSS. StarRocks delivers three capabilities: scalar filtering (spatiotemporal/weather/tags/fault conditions), vector search (whole-image and crop similar scenes), hybrid retrieval (full-text + vector multi-recall with RRF fusion). This supports workflows like "input fault image vector + rainy night condition → retrieve similar driving frames → output video clips and hit timestamps" for model training supplementation and fault traceback.

StarRocks hybrid retrieval
StarRocks hybrid retrieval

Flink: From Stream Processing to Streaming Agent Runtime

Flink undergoes the largest paradigm shift: from stream data processing engine to enterprise-grade Streaming Agent runtime platform, building Agentic Lake via lake-stream unification.

Traditional streaming processes "data flows"; Streaming Agents process "event flows": data arrives continuously, Agent state evolves incrementally per event, triggering LLM inference at key moments, outputting structured results and business actions.

Three-layer Streaming Agent architecture:

Base layer: Streaming Pipeline, state memory management, adaptive elastic scaling, end-to-end fault self-healing.

Operator layer: Multi-source ingestion (media streams/log streams/storage systems), media processing (decoding/frame extraction/slicing), perception understanding (ASR/OCR/VLM/object detection), retrieval generation (RAG/LLM/TTS), Agent operators (Workflow/ReAct, MCP/Tools extension).

Business layer: Reusable industry solutions.

Flink integrates Alibaba Cloud Bailian inference service, supporting mixed calls of specialized perception models and general LLMs. Agent global context evolves continuously across events and modalities.

Real-time sports commentary Agent — powered CCTV Asian Games broadcast. Six-step pipeline: event signal ingest → video preprocessing (format standardization, key-frame sampling) → event observation (independent plugins extract on-screen actions, score graphics, person targets) → event understanding (elements combined into semantics, key events and progress identified) → commentary planning (key-event priority, factual reliability judgment) → broadcast generation (speech generation, digital verification, subtitle TTS real-time playout). Flink's built-in AI service uniformly hosts CV/OCR/VLM vision understanding and TTS speech generation; continuous Flink execution guarantees zero state loss.

Store intelligent operations Agent — real-time business event detection from camera streams. Adapts to checkout, staff restocking, empty shelves, in-store disputes. Perception extracts only visible facts without direct judgment, preserving counter-evidence and evidence frames. Understanding maintains cross-window state per camera; silent or ending actions auto-close events. Planning uses evidence-sufficiency tiered review — rule confirmation → re-check — balancing cost and accuracy. Execution writes structured events and evidence frames to Paimon lake tables, powering dashboards, alerts, and retrospectives for a fully traceable data chain.

Flink Streaming Agent architecture
Flink Streaming Agent architecture

Open Source, Open Platform: Infrastructure for AI-Native Applications

As Agents become new data platform users and multimodal becomes the dominant data form, infrastructure must be fully reconstructed. DLF unifies multimodal management with Paimon 2.0 + Fluss lake-stream integration; EMR Ray + Daft connects data processing to model training; StarRocks enables precise hybrid retrieval for Agents; Flink drives real-time business actions via Streaming Agents — all coordinated with Milvus vector engine to form the complete Agentic Lake + Agentic Streaming stack. Alibaba Cloud's open big data platform offers fully managed services with deep optimizations, giving enterprises open-source flexibility plus enterprise-grade stability and performance. From "lake births all things" to "empowering AI," Alibaba Cloud's open big data platform is becoming the inevitable choice for building Agent-era multimodal data infrastructure.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

FlinkStarRocksAlibaba CloudMultimodal DataDLFAgentic StreamingAgentic LakeEMR Ray
Alibaba Cloud Big Data AI Platform
Written by

Alibaba Cloud Big Data AI Platform

The Alibaba Cloud Big Data AI Platform builds on Alibaba’s leading cloud infrastructure, big‑data and AI engineering capabilities, scenario algorithms, and extensive industry experience to offer enterprises and developers a one‑stop, cloud‑native big‑data and AI capability suite. It boosts AI development efficiency, enables large‑scale AI deployment across industries, and drives business value.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.