Recap of Tencent Cloud AI DLC Launch: Serverless Spark + Ray Integration for Agent‑Native Data Lake
The article reviews Tencent Cloud's AI DLC launch, detailing how the serverless Spark + Ray platform unifies data, compute, and agent workflows, introduces four architectural upgrades, showcases core engines (TCRay, Xpark, Meson, Open Engine), and presents benchmark results and real‑world practices from Bosch and WorkBuddy that demonstrate significant performance and productivity gains.
The DataFun Agentic AI Summit in Shenzhen featured the official release of Tencent Cloud's Intelligent Data Lake Computing (AI DLC). The product is positioned as the industry's first production‑grade serverless Spark + Ray fully managed platform, aiming to provide a unified infrastructure where data, compute, models, agent trajectories, and feedback continuously flow.
Design Principle : In the Agent era, the data foundation must no longer treat data processing, model training, online inference, and agent execution as isolated systems. Instead, it should enable data, compute, models, trajectories, and feedback to move together on a single platform, delivering trustworthy, explainable, and instantly callable context for each agent decision.
Four Upgrades introduced by AI DLC address the four dimensions of the AI/Agent shift: (1) data evolving from structured to multimodal, (2) resources moving from CPU‑only to heterogeneous CPU + GPU, (3) training shifting from pre‑training to private‑data post‑training, and (4) data engineering transitioning from deterministic one‑time pipelines to iterative, trajectory‑driven systems. These upgrades are realized through four design ideas: Ray × Spark integration, Xpark + Ray Libraries + Meson, TCLake for unified metadata and lineage, and an Open Engine ecosystem (MCP, Skills, SDK, Open API) that supports both internal and external agents.
Core Engines :
TCRay converts the open‑source Ray runtime into a production‑grade platform. It has been stress‑tested with 3,000 nodes and 100,000 actors, achieving a 40% reduction in GCS steady‑state memory and a 90% drop in idle‑worker memory. In a 12‑hour workload with 13,390 events, History Server load time fell from minutes to seconds (‑98%), and peak memory usage decreased by 99%.
Xpark focuses on multimodal data processing. In a CLIP image‑text matching test, Xpark's inference throughput was three times higher than the open‑source Data‑Juicer, with GPU utilization consistently near 100%. Its Exact Substring Dedup operator outperformed the open‑source Text‑dedup single‑node version by 47.8× and reduced large‑model request volume by 70% through model cascading and discriminative operators.
Meson is a high‑performance vectorized engine for structured data. In a 1 TB TPC‑DS benchmark, Meson delivered 3.6× overall speedup over community Spark and up to 5× on compute‑intensive queries, while CPU usage dropped from 80% to 40% and I/O throughput rose from 12 Gbps to 20 Gbps. Compatibility with Spark API, SQL, UDFs, and ecosystem components eases migration.
Open Engine (including TCLake, MCP, Skills, SDK, Open API) provides unified metadata, data‑plus‑AI lineage, and a programmable interface that allows both internal and external agents to orchestrate resources, jobs, and results.
Real‑World Practices :
Bosch China demonstrated hybrid scheduling with Ray on TKE, integrating Ray Job, Ray Serve, Spark, and custom Pods into a unified pool. Their key benefits were "development becomes production line" and "queueing with pre‑emptive scheduling", eliminating resource islands and enabling consistent environments across development, training, and inference.
WorkBuddy showcased an end‑to‑end data pipeline where business data, dialogue documents, and logs flow through object storage, TCLake, Serverless Spark × Ray, and Xpark for multimodal processing, post‑training, and inference. Columnar reads, hot‑cold separation, and Xpark multimodal operators reduced unnecessary scans, while API‑driven agent orchestration linked lake trajectories, feature extraction, Ray Train, and evaluation. Reported gains include a reduction of data‑processing time from 5 hours to 1 hour (5× throughput), 80% lower CPU usage, a 70% shorter delivery cycle, and a 90% decrease in manual hand‑offs.
Overall, the AI DLC release and the showcased practices illustrate a shift from traditional ETL‑centric platforms to an AI‑native data lake that continuously cycles data, compute, models, and agent feedback under unified governance. Tencent Cloud plans to continue evolving the platform in 2026 with a focus on "more, faster, better, cheaper" capabilities, expanding product iterations, co‑building with customers, and fostering an open ecosystem.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DataFunTalk
Dedicated to sharing and discussing big data and AI technology applications, aiming to empower a million data scientists. Regularly hosts live tech talks and curates articles on big data, recommendation/search algorithms, advertising algorithms, NLP, intelligent risk control, autonomous driving, and machine learning/deep learning.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
