Tagged articles

Data Lake

388 articles · Page 1 of 4
ITPUB
ITPUB
Sep 14, 2026 · Big Data

Flink CDC vs Canal: Building 99.99% Consistent Real-Time Data Pipelines

This article compares Flink CDC with traditional Canal-based architectures, explains Flink CDC's core principles including lock-free snapshots and exactly-once semantics, provides DataStream and SQL code examples for MySQL integration, and covers five high-frequency interview questions on DDL handling, DELETE capture, and large-table optimization.

Data LakeDebeziumExactly-Once
0 likes · 12 min read
Flink CDC vs Canal: Building 99.99% Consistent Real-Time Data Pipelines
Smart Sea Tide
Smart Sea Tide
Sep 7, 2026 · Databases

Data Governance: 12 Brutal Truths Behind Common Excuses

This article dismantles twelve common data governance fallacies — from demanding instant metadata completion to treating governance as a cost center — arguing that effective governance requires confronting upstream data pollution, aligning incentives across departments, and prioritizing critical data paths over blanket coverage.

Data LakeData Qualitydata governance
0 likes · 10 min read
Data Governance: 12 Brutal Truths Behind Common Excuses
ITPUB
ITPUB
Aug 20, 2026 · Industry Insights

2026 China Database Technology Conference Launches: Data Fusion and AI Leadership

The 17th China Database Technology Conference (DTCC 2026) ran from August 20‑22 in Beijing, gathering top experts to discuss database kernel innovations, cloud‑native and distributed practices, AI‑driven data, vector databases, real‑time warehouses, and the emerging Agent era, while showcasing cutting‑edge solutions from Dameng, Tencent Cloud, Alibaba Cloud, OceanBase and GoldenDB.

AIAgentData Lake
0 likes · 15 min read
2026 China Database Technology Conference Launches: Data Fusion and AI Leadership
Smart Sea Tide
Smart Sea Tide
Aug 18, 2026 · Big Data

How to Choose and Architect a Data Lake Platform for Enterprise Digital Transformation

The article outlines the strategic need for a unified data lake in a digital‑focused enterprise, details functional and non‑functional requirements such as linear scalability, real‑time and batch processing, multi‑tenant support, security and governance, and presents a comprehensive architecture design that integrates storage, compute, and management components.

Data Lakebig datacloud-native
0 likes · 19 min read
How to Choose and Architect a Data Lake Platform for Enterprise Digital Transformation
DataFunSummit
DataFunSummit
Aug 17, 2026 · Industry Insights

AI Era Data Infrastructure: From Storing Data to Enabling Agent‑Driven Context

The article analyzes how the rise of AI agents transforms data platforms from simple storage and query engines into AI‑native systems that provide trustworthy, real‑time context for autonomous decision‑making, outlining the three‑layer evolution of storage, compute, and application and the architectural upgrades required for modern data lakes.

AIAgentCloud Computing
0 likes · 13 min read
AI Era Data Infrastructure: From Storing Data to Enabling Agent‑Driven Context
Big Data Technology Tribe
Big Data Technology Tribe
Aug 14, 2026 · Big Data

How lance‑spark Implements Blob V2 Support: A Deep Dive

The article explains how lance‑spark enables Lance's Blob V2 storage by marking a column with blob encoding and setting file_format_version ≥ 2.2, describes the metadata‑driven descriptor schema, the write path that still accepts Spark BINARY, and the size‑based placement of binary data into inline, packed, dedicated or external blob files.

ArrowBlob V2Data Lake
0 likes · 12 min read
How lance‑spark Implements Blob V2 Support: A Deep Dive
Data Integration and Governance
Data Integration and Governance
Aug 11, 2026 · Big Data

What Is Data Architecture? Clarifying Databases, Data Warehouses, Lakes, and Middle Platforms

As enterprises add more systems, data volumes explode while accessing it becomes harder; this article defines data architecture, explains how databases, data warehouses, data lakes, and data middle platforms each solve distinct layers, and outlines the key questions and challenges for building a unified, reusable data ecosystem.

Data Lakebig datadata architecture
0 likes · 13 min read
What Is Data Architecture? Clarifying Databases, Data Warehouses, Lakes, and Middle Platforms
DataFunTalk
DataFunTalk
Aug 4, 2026 · Big Data

Recap of Tencent Cloud AI DLC Launch: Serverless Spark + Ray Integration for Agent‑Native Data Lake

The article reviews Tencent Cloud's AI DLC launch, detailing how the serverless Spark + Ray platform unifies data, compute, and agent workflows, introduces four architectural upgrades, showcases core engines (TCRay, Xpark, Meson, Open Engine), and presents benchmark results and real‑world practices from Bosch and WorkBuddy that demonstrate significant performance and productivity gains.

AI DLCBenchmarkData Lake
0 likes · 14 min read
Recap of Tencent Cloud AI DLC Launch: Serverless Spark + Ray Integration for Agent‑Native Data Lake
Xike
Xike
Aug 1, 2026 · Big Data

Big Data Series #6: Getting Started with Iceberg Lakehouse – Snapshots, Time Travel, and Writes

This tutorial explains how Apache Iceberg adds snapshot metadata to files on HDFS or object storage, enabling atomic writes, time‑travel queries, and schema evolution, and walks through a Docker‑based setup with a REST catalog, Spark 3.5.3, and hands‑on examples including table creation, data insertion, version queries, and troubleshooting tips.

Apache IcebergData LakeHDFS
0 likes · 16 min read
Big Data Series #6: Getting Started with Iceberg Lakehouse – Snapshots, Time Travel, and Writes
YiSu Grain
YiSu Grain
Jul 28, 2026 · Big Data

Data Warehouse vs Data Lake vs Lambda/Kappa: Choosing the Right Architecture

This article explains why online transaction databases (OLTP) should not be used for multi‑year analytics, outlines the four core characteristics of a data warehouse, details the ETL process, compares star and snowflake schemas, contrasts data warehouses with data lakes, and guides you through selecting Lambda or Kappa architectures using a concrete retail‑store scenario.

Data LakeETLKappa Architecture
0 likes · 36 min read
Data Warehouse vs Data Lake vs Lambda/Kappa: Choosing the Right Architecture
Data Integration and Governance
Data Integration and Governance
Jul 22, 2026 · Big Data

Data Warehouse vs Data Mart vs Data Lake vs Data Middle Platform: Clear Differences

The article explains how data warehouses, data marts, data lakes, and data middle platforms each address distinct problems—unified analytics, departmental needs, raw data storage, and governed reusable capabilities—while outlining their relationships, typical use cases, implementation considerations, and guidance on choosing the right architecture for a given business stage.

Data LakeData Martbig data
0 likes · 16 min read
Data Warehouse vs Data Mart vs Data Lake vs Data Middle Platform: Clear Differences
DataFunSummit
DataFunSummit
Jul 22, 2026 · Big Data

How Tencent Redefines Data Architecture for the Agent Era

With agents moving from Q&A to execution, traditional architectures expose three critical flaws—data stored in lakes, models in the cloud, and split scheduling—forcing petabyte‑scale data movement; Tencent Cloud’s big data AI DLC resolves this by running Spark and Ray side‑by‑side on the same lake, enabling closed‑loop processing and automatic trajectory capture.

AIAgentData Lake
0 likes · 2 min read
How Tencent Redefines Data Architecture for the Agent Era
DataFunTalk
DataFunTalk
Jul 18, 2026 · Big Data

How Tencent Redefines Data Architecture for the Agent Era

The article analyzes how traditional data‑lake‑model‑cloud architectures expose three critical flaws for agentic AI—massive data movement, fragmented logging, and split compute—then details Tencent Cloud's Big Data AI DLC solution that unifies Spark and Ray on a single lake to enable in‑place processing, closed‑loop training, and cost‑effective iteration.

Agentic AIData LakeRay
0 likes · 2 min read
How Tencent Redefines Data Architecture for the Agent Era
DataFunSummit
DataFunSummit
Jul 11, 2026 · Artificial Intelligence

Tencent CodeBuddy’s AI DLC Slashes Training Time and Costs with a Unified Spark‑Ray Service

The article explains how Tencent CodeBuddy’s AI DLC platform unifies Spark batch processing and Ray training to eliminate data movement, turning agent trajectories into reusable training fuel, which reduces monthly‑level training cycles to weekly, enables in‑place computation on billions of features, and cuts operational costs by 60%.

AI DLCData LakeGPU utilization
0 likes · 2 min read
Tencent CodeBuddy’s AI DLC Slashes Training Time and Costs with a Unified Spark‑Ray Service
Data Integration and Governance
Data Integration and Governance
Jul 9, 2026 · Big Data

Data Warehouse vs Big Data Platform vs Data Lake vs Data Middle Platform vs Lake‑Warehouse Integration: What’s the Real Difference?

The article compares five data‑architecture concepts—data warehouse, big data platform, data lake, data middle platform, and lake‑warehouse integration—explaining the specific problems each solves, their core characteristics, advantages, risks, and guidance on when to adopt each solution.

Data IntegrationData LakeLake‑Warehouse Integration
0 likes · 12 min read
Data Warehouse vs Big Data Platform vs Data Lake vs Data Middle Platform vs Lake‑Warehouse Integration: What’s the Real Difference?
ITPUB
ITPUB
Jul 2, 2026 · Industry Insights

How ColdFront Sets pgEdge Apart in the OLTP‑OLAP‑AI Showdown

The article compares four emerging data‑lake‑for‑PostgreSQL solutions—Databricks LTAP, EDB Fusion Analytics, Snowflake pg_lake, and pgEdge's ColdFront—highlighting ColdFront's unique transparent Iceberg layer, writable cold data, DuckDB integration, and the strategic trade‑offs developers must weigh when choosing a modern OLTP/OLAP/AI architecture.

Agentic AIColdFrontData Lake
0 likes · 9 min read
How ColdFront Sets pgEdge Apart in the OLTP‑OLAP‑AI Showdown
DeepNoMind
DeepNoMind
Jun 27, 2026 · Backend Development

Designing a Production‑Grade Distributed Logging and Metrics Platform

This article presents an end‑to‑end design of a production‑grade observability platform that ingests millions of real‑time logs, metrics, and events, detailing functional and non‑functional requirements, capacity planning, component choices such as Kafka, Flink, Elasticsearch, object‑storage data lakes, and the trade‑offs involved.

Data LakeElasticsearchFlink
0 likes · 21 min read
Designing a Production‑Grade Distributed Logging and Metrics Platform
Alibaba Cloud Native
Alibaba Cloud Native
Jun 26, 2026 · Cloud Native

One-Click Real-Time Stream Ingestion: Alibaba Cloud Kafka’s Native Data Lake Integration

Alibaba Cloud Message Queue for Kafka introduces a native message‑to‑lake capability that integrates Apache Iceberg with OSS Table Bucket, eliminating Spark/Flink/Kafka Connect, providing exactly‑once semantics, automatic schema management, dual write modes, smart partitioning, and up to ten‑fold performance gains across diverse real‑time analytics scenarios.

Apache IcebergData LakeExactly-Once
0 likes · 12 min read
One-Click Real-Time Stream Ingestion: Alibaba Cloud Kafka’s Native Data Lake Integration
ByteDance Data Platform
ByteDance Data Platform
Jun 24, 2026 · Artificial Intelligence

How AI Is Redefining Data Products: New Paths for Enterprise Intelligence

The article analyzes how the AI era shifts data from a passive by‑product to a core driver of large‑model performance, traces the evolution of data products from the DBA era through big‑data to AI‑native solutions, and details Volcano Engine’s four‑layer AI data platform that closes the data‑to‑model‑to‑Agent loop.

AIAgentData Lake
0 likes · 12 min read
How AI Is Redefining Data Products: New Paths for Enterprise Intelligence
Alibaba Cloud Native
Alibaba Cloud Native
Jun 19, 2026 · Big Data

Why Real-Time Data Lake Ingestion Is Dropping ETL in the AI Era: Architecture Simplification from Kafka to Iceberg

In the AI‑driven era, enterprises need a data foundation that supports both real‑time consumption and long‑term historical analysis, and the emerging "zero‑ETL" trend moves generic ingestion capabilities from external Flink/Spark jobs into a streamlined Kafka‑to‑Iceberg pipeline, reducing complexity while preserving low latency, consistency, schema evolution, CDC semantics and open‑ecosystem compatibility.

Data LakeKafkaStreaming
0 likes · 25 min read
Why Real-Time Data Lake Ingestion Is Dropping ETL in the AI Era: Architecture Simplification from Kafka to Iceberg
Alibaba Cloud Developer
Alibaba Cloud Developer
Jun 18, 2026 · Big Data

How AI-Driven Real-Time Data Lakes Are Ditching ETL: A Kafka‑to‑Iceberg Architecture Simplification

In the AI era, enterprises need a data foundation that supports both low‑latency streaming and long‑term analytics, and the combination of Kafka, Iceberg and object storage is emerging as a preferred solution; by moving ingestion capabilities closer to the message layer and eliminating external ETL jobs, a "zero‑ETL" approach reduces architectural complexity, improves consistency, and streamlines schema evolution and small‑file management.

CDCData LakeKafka
0 likes · 27 min read
How AI-Driven Real-Time Data Lakes Are Ditching ETL: A Kafka‑to‑Iceberg Architecture Simplification
StarRocks
StarRocks
Jun 4, 2026 · Databases

How StarRocks and Iceberg Enable Federated Queries: A Practical Walkthrough

This article details Fresha's real‑world integration of StarRocks with Apache Iceberg, covering metadata planning, distributed execution, adaptive metadata retrieval, hot‑cold data layering, missing statistics handling, catalog configuration, and performance optimizations that together demonstrate how federated queries can be efficiently executed over data‑lake tables.

Apache IcebergData LakeFederated Query
0 likes · 14 min read
How StarRocks and Iceberg Enable Federated Queries: A Practical Walkthrough
DataFunSummit
DataFunSummit
May 25, 2026 · Big Data

How Hisense Built an AI‑Ready Multimodal Data Platform: Storage, Governance, and Development

This article details Hisense's journey to create an AI‑ready multimodal data platform, covering the challenges of integrating diverse business systems, the shift from a Hadoop‑based architecture to a cloud‑native data lake, the JuData governance and development platform, and six practical scenarios that demonstrate unified ingestion, metadata management, rule‑based quality control, intelligent asset retrieval, and future AI‑driven DataOps capabilities.

AI platformData LakeDataOps
0 likes · 23 min read
How Hisense Built an AI‑Ready Multimodal Data Platform: Storage, Governance, and Development
DataFunSummit
DataFunSummit
May 21, 2026 · Big Data

Alibaba Cloud’s Agent-Ready Big Data AI Infrastructure: Boosting Data Development from Hours to Minutes

Facing a projected 85% of enterprises deploying internal agents within two years, Alibaba Cloud proposes an Agent-Ready big‑data AI infrastructure—comprising a unified data lake, real‑time processing, high‑dimensional vector retrieval, elastic model serving, and comprehensive security governance—that has already cut data‑development cycles from hours to 5‑10 minutes in internal model‑training and Taobao flash‑sale scenarios.

AIAgent-ReadyData Lake
0 likes · 15 min read
Alibaba Cloud’s Agent-Ready Big Data AI Infrastructure: Boosting Data Development from Hours to Minutes
DataFunSummit
DataFunSummit
May 20, 2026 · Big Data

How Kuaishou’s Real‑Time Data Lake Boosts AI and BI Architecture

The article explains how Kuaishou partnered with Apache Hudi to overhaul its ODS‑based data lake, addressing latency, storage cost, and complexity for AI and BI workloads, detailing the evolution from mysql‑to‑hive to mysql‑to‑hudi 1.0 and 2.0, the resulting performance gains, cost savings, and future roadmap.

AIBIData Lake
0 likes · 20 min read
How Kuaishou’s Real‑Time Data Lake Boosts AI and BI Architecture
DataFunSummit
DataFunSummit
May 11, 2026 · Artificial Intelligence

How Lance Powers Enterprise Multimodal AI Data Lakes

The article analyzes why 74% of AI projects fail due to feedback gaps and data silos, explains how the open‑source Lance format addresses these issues with unified multimodal storage, outlines a layered Lance‑on‑Ray architecture, and details three real‑world practices—implicit feedback loops, GPU‑accelerated self‑evolution, and semantic knowledge‑graph evolution—to boost R&D efficiency.

CAGRADaftData Lake
0 likes · 13 min read
How Lance Powers Enterprise Multimodal AI Data Lakes
DataFunSummit
DataFunSummit
May 5, 2026 · Big Data

A New Data Lake Paradigm: Volcano Engine’s Multi‑Modal Data Lake Built on Lance

The article presents Volcano Engine’s AI‑focused data lake built on the Lance format, detailing why traditional lakes fall short for multimodal data, the engineering enhancements such as Binary Copy Compaction, Lance Insight, distributed vector indexing, JSON‑based tagging, Row‑ID shuffle optimization, and real‑world case studies that demonstrate significant performance and cost gains.

AIBinary Copy CompactionData Lake
0 likes · 18 min read
A New Data Lake Paradigm: Volcano Engine’s Multi‑Modal Data Lake Built on Lance
Smart Sea Tide
Smart Sea Tide
Apr 29, 2026 · Cloud Computing

Data as a Service (DaaS): Architecture and Key Advantages

The article explains how Data as a Service (DaaS) builds on data lakes and SaaS models to centralize governance, cut duplication and infrastructure costs, accelerate real‑time analytics, support cloud‑native deployments, and enable mobile/web applications through unified APIs.

DaaSData LakeData as a Service
0 likes · 8 min read
Data as a Service (DaaS): Architecture and Key Advantages
Data Integration and Governance
Data Integration and Governance
Apr 28, 2026 · Big Data

Data Warehouse, Data Lake, Data Middle Platform, Lake‑Warehouse Integration: Key Differences

This article systematically explains the definitions, architectures, advantages, disadvantages, and evolution of data warehouses, big‑data platforms, data lakes, data middle platforms, and lake‑warehouse integration, helping practitioners choose the right terminology and technology for their projects.

Data LakeFineDataLinkLake‑Warehouse Integration
0 likes · 12 min read
Data Warehouse, Data Lake, Data Middle Platform, Lake‑Warehouse Integration: Key Differences
DataFunSummit
DataFunSummit
Apr 19, 2026 · Big Data

How OPPO Built a Multi‑Modal Data Lake with Gravitino and Curvine

OPPO’s data‑lake team, led by David, detailed their transition from Hive‑Spark to a unified multi‑modal lake, leveraging Gravitino for cross‑engine metadata management and the open‑source Curvine cache to eliminate data silos, boost I/O performance, and support massive image, recommendation, and AI‑Agent workloads.

Data LakeMultimodalbig data
0 likes · 11 min read
How OPPO Built a Multi‑Modal Data Lake with Gravitino and Curvine
DataFunTalk
DataFunTalk
Apr 18, 2026 · Databases

How Will Apache Doris Evolve in 2026 to Power AI‑Driven Data Workloads?

The article outlines Apache Doris's 2026 roadmap, detailing how the database will shift from pure analytics to a unified AI‑enabled platform with enhanced semi‑structured data support, vector and hybrid search, agent‑focused capabilities, and expanded storage and lakehouse integrations to meet emerging AI workloads.

AI IntegrationApache DorisData Lake
0 likes · 14 min read
How Will Apache Doris Evolve in 2026 to Power AI‑Driven Data Workloads?
Past Memory Big Data
Past Memory Big Data
Mar 27, 2026 · Big Data

Why AI Workloads Require Rebuilding Parquet: A Deep Dive into Lance

The article explains how traditional Parquet‑based lakehouse architectures, optimized for large‑scale scans, struggle with AI workloads that need ultra‑low‑latency random access, and how Lance redesigns the storage format, indexing and write path to provide O(1) addressing, native vector support, and seamless integration with native execution engines.

AI workloadsData LakeLance
0 likes · 12 min read
Why AI Workloads Require Rebuilding Parquet: A Deep Dive into Lance
DataFunTalk
DataFunTalk
Mar 3, 2026 · Big Data

Exploring Tencent Cloud’s Iceberg Batch‑Stream Integration and AI‑Driven Data Governance

This article presents a series of seven technical case studies—including Tencent Cloud’s Iceberg‑based batch‑stream integration, AI‑driven data governance with Apache Gravitino, Xiaohongshu’s lakehouse evolution, and a multimodal data‑lake solution—detailing challenges, architectural designs, implementation steps, performance results, and future directions.

AIData LakeMultimodal
0 likes · 8 min read
Exploring Tencent Cloud’s Iceberg Batch‑Stream Integration and AI‑Driven Data Governance
StarRocks
StarRocks
Feb 11, 2026 · Big Data

How StarRocks and Apache Paimon Build a True Lakehouse Native Engine

This article details the deep integration of StarRocks with Apache Paimon, describing the unified architecture, version evolution, performance enhancements, time‑travel queries, native readers/writers, distributed planning, and future roadmap for achieving lakehouse‑native analytics at scale.

Apache PaimonData LakeLakehouse
0 likes · 10 min read
How StarRocks and Apache Paimon Build a True Lakehouse Native Engine
DataFunSummit
DataFunSummit
Feb 8, 2026 · Big Data

Kuaishou’s Data Lake Upgrade with Hudi: Solving AI & BI Challenges

The article explains how Kuaishou modernized its data lake by partnering with Apache Hudi to address latency, storage cost, and consistency issues in both AI and BI pipelines, detailing architectural changes, new ingestion tools, partitioning strategies, compaction mechanisms, performance gains and future plans.

AIBIData Lake
0 likes · 20 min read
Kuaishou’s Data Lake Upgrade with Hudi: Solving AI & BI Challenges
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Feb 4, 2026 · Big Data

How Paimon + StarRocks Power Real‑Time OLAP for Double‑11 Mega‑Sales

During Double‑11 mega‑sales, Taobao Group faced exploding OLAP query traffic, costly data sync pipelines, and slow near‑real‑time analytics, so they unified real‑time and batch data in Paimon, leveraged StarRocks for high‑performance lake queries, tuned cluster settings, and saved nearly ten‑million yuan annually while cutting refresh latency by 80%.

Data LakeOLAPPaimon
0 likes · 22 min read
How Paimon + StarRocks Power Real‑Time OLAP for Double‑11 Mega‑Sales
Lakehouse Research Base
Lakehouse Research Base
Jan 29, 2026 · Big Data

Unstructured Data Lake Ingestion: File Body + Metadata Registration Patterns

This article details the industry-standard dual pattern of file body storage and metadata registration for unstructured data lake ingestion, covering a three-layer architecture, batch and real-time implementation steps using Paimon and Flink, AI-specific optimizations like label standardization and format conversion, and key operational considerations for permissions, versioning, and cost control.

AI data preparationData LakeFlink
0 likes · 11 min read
Unstructured Data Lake Ingestion: File Body + Metadata Registration Patterns
Data Integration and Governance
Data Integration and Governance
Jan 14, 2026 · Big Data

Finally, a Clear Guide to Data Architecture

This article explains data architecture from the ground up, covering data sources, storage options such as operational databases, data warehouses and data lakes, ETL processing steps, layered data modeling, service delivery methods, and governance practices to ensure reliable, secure, and business‑driven data management.

Data LakeETLdata architecture
0 likes · 11 min read
Finally, a Clear Guide to Data Architecture
Lakehouse Research Base
Lakehouse Research Base
Jan 8, 2026 · Big Data

SME Data Architecture Selection: Lightweight Lakehouse Implementation Guide

This article provides a practical framework for SMEs to select and implement data platform architectures, comparing traditional warehouses, lakehouse, and multimodal data lakes across cost, ROI, and operational efficiency, with scenario-based recommendations and a lightweight lakehouse implementation guide using open-source stack Flink, Paimon, StarRocks, and MinIO.

Data LakeFlinkLakehouse
0 likes · 21 min read
SME Data Architecture Selection: Lightweight Lakehouse Implementation Guide
Big Data Tech Team
Big Data Tech Team
Dec 29, 2025 · Big Data

Data Warehouse vs Data Mart vs Data Lake: Which Should Your Enterprise Choose?

The article explains the distinct roles of data warehouses, data marts, and data lakes, illustrates their differences with analogies and real‑world cases, outlines a three‑step strategy for enterprises, highlights common pitfalls, and offers a decision guide to help organizations choose the right architecture for their data needs.

Data LakeData Martdata warehouse
0 likes · 11 min read
Data Warehouse vs Data Mart vs Data Lake: Which Should Your Enterprise Choose?
DataFunTalk
DataFunTalk
Dec 26, 2025 · Cloud Native

How Haier Built a Cloud‑Native Multi‑Modal Data Lake for AI‑Ready Manufacturing

Haier’s digital transformation leverages a cloud‑native, open‑source‑based multi‑modal data lake that unifies structured and unstructured industrial data, uses metadata models and knowledge graphs for governance, and provides AI‑ready services that balance performance, cost, and real‑time requirements.

AIData LakeMultimodal Data
0 likes · 12 min read
How Haier Built a Cloud‑Native Multi‑Modal Data Lake for AI‑Ready Manufacturing
JD Tech Talk
JD Tech Talk
Dec 12, 2025 · Big Data

Understanding Hudi Core Concepts: Timeline, Indexes, and Table Types Explained

This article explains Apache Hudi’s core concepts, including its timeline architecture, file layout, indexing mechanisms, and the two primary table types—Copy on Write and Merge on Read—along with their trade‑offs and the various query modes such as snapshot, time‑travel, and incremental queries.

Apache HudiData LakeIndexing
0 likes · 9 min read
Understanding Hudi Core Concepts: Timeline, Indexes, and Table Types Explained
JD Cloud Developers
JD Cloud Developers
Dec 12, 2025 · Big Data

Apache Hudi Core Concepts: Timeline, Indexes, Table Types & Queries

This article explains Apache Hudi’s core architecture, detailing the timeline mechanism, file layout, indexing strategies, the two main table types (Copy‑On‑Write and Merge‑On‑Read), and various query modes such as snapshot, time‑travel, read‑optimized and incremental queries.

Apache HudiData LakeTable Types
0 likes · 9 min read
Apache Hudi Core Concepts: Timeline, Indexes, Table Types & Queries
Past Memory Big Data
Past Memory Big Data
Dec 12, 2025 · Big Data

How Uber Reduced Data Freshness from Hours to Minutes Using Flink Streaming

Uber rebuilt its data‑lake ingestion pipeline with Apache Flink, replacing batch jobs with a streaming architecture that cuts data freshness from hours to minutes, lowers compute usage by 25%, and solves challenges like small‑file proliferation, partition skew, and checkpoint‑commit synchronization at petabyte scale.

Apache FlinkApache HudiData Freshness
0 likes · 10 min read
How Uber Reduced Data Freshness from Hours to Minutes Using Flink Streaming
DataFunSummit
DataFunSummit
Dec 1, 2025 · Big Data

7 Cutting-Edge Data Engineering Practices Shaping AI-Driven Data Lakes

This article collection showcases seven advanced data engineering solutions—from Tencent Cloud's Iceberg batch‑stream integration and Apache Gravitino metadata lineage to Xiaohongshu's Lakehouse evolution and multimodal AI data lake implementations—highlighting architectural innovations, performance optimizations, and real‑world deployment insights for modern big‑data platforms.

Apache GravitinoApache IcebergBatch-Stream Integration
0 likes · 7 min read
7 Cutting-Edge Data Engineering Practices Shaping AI-Driven Data Lakes
Past Memory Big Data
Past Memory Big Data
Dec 1, 2025 · Big Data

Apache XTable: A Universal Translator for Data Lake Format Interoperability

Apache XTable introduces a lightweight metadata translation layer that decouples data storage from format metadata, enabling zero‑copy, omni‑directional conversion among Hudi, Iceberg, and Delta Lake, allowing organizations to write with one format and read with any engine without duplicating Parquet files.

Apache XTableData LakeDelta Lake
0 likes · 7 min read
Apache XTable: A Universal Translator for Data Lake Format Interoperability
DataFunSummit
DataFunSummit
Nov 24, 2025 · Big Data

How Tencent Cloud Uses Iceberg, Gravitino and Multimodal Lakes for Unified Data Processing

This article series explores Tencent Cloud's Iceberg‑based batch‑stream integration, Apache Gravitino's unified metadata and lineage solution, Xiaohongshu's data‑architecture evolution for the Big AI Data era, and a practical Data+AI multimodal data‑lake implementation, highlighting challenges, architectural designs, and performance gains.

Data LakeMultimodal Databig data
0 likes · 7 min read
How Tencent Cloud Uses Iceberg, Gravitino and Multimodal Lakes for Unified Data Processing
DataFunTalk
DataFunTalk
Nov 22, 2025 · Big Data

How Modern Data Lakes and AI Governance Transform Enterprise Analytics

This article collection examines Tencent Cloud’s Iceberg batch‑stream integration, AI‑driven game data governance, Apache Gravitino unified metadata and lineage, Xiaohongshu’s multimodal data‑lake evolution, and Volcano Engine’s Data+AI multimodal lake, highlighting architectures, techniques, performance gains, and practical implementations.

AI GovernanceData LakeGravitino
0 likes · 7 min read
How Modern Data Lakes and AI Governance Transform Enterprise Analytics
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Nov 15, 2025 · Big Data

From a Decade-Long Big Data Journey to a Cloud‑Native Lakehouse

This article chronicles a ten‑year evolution of a self‑built big data platform—detailing early Hadoop clusters, successive migrations to Spark, Hive, Hudi, and StarRocks, the operational challenges encountered, and the comprehensive shift to Alibaba Cloud EMR Serverless that delivered significant cost, performance, and stability gains while outlining future intelligent‑ecosystem plans.

Data LakeEMR ServerlessSpark
0 likes · 17 min read
From a Decade-Long Big Data Journey to a Cloud‑Native Lakehouse
Smart Sea Tide
Smart Sea Tide
Nov 4, 2025 · Big Data

Implementing an Integrated Data Lake and Lakehouse Architecture with Apache Iceberg and Flink

The article explains the concepts of data lakes and lakehouses, compares them with traditional data warehouses, outlines the reliability, performance, and security challenges of data lakes, and then details a practical lakehouse implementation using Apache Iceberg, Flink SQL, CDC pipelines, and supporting tools such as Hive Metastore and Trino.

Apache IcebergCDCData Lake
0 likes · 18 min read
Implementing an Integrated Data Lake and Lakehouse Architecture with Apache Iceberg and Flink
DataFunSummit
DataFunSummit
Oct 6, 2025 · Artificial Intelligence

Why Vector Lakes Are the Next Frontier for AI Data Management

This article explains how Zilliz's Vector Lake extends traditional data lakes with a unified storage‑compute architecture optimized for massive unstructured and vector data, detailing its background, key data types, autonomous‑driving use case, data flow, architecture, and deployment options.

AI Data ManagementData LakeVector Lake
0 likes · 13 min read
Why Vector Lakes Are the Next Frontier for AI Data Management
IT Architects Alliance
IT Architects Alliance
Sep 21, 2025 · Big Data

From Data Warehouses to Lakehouses: Why Data Architecture Keeps Evolving

This article traces the three‑generation evolution of data architecture—from the structured‑data era of data warehouses, through the flexible, multi‑format data lake, to the unified lakehouse model—explaining the drivers, benefits, challenges, and future trends shaping modern data platforms.

Data LakeLakehousedata architecture
0 likes · 11 min read
From Data Warehouses to Lakehouses: Why Data Architecture Keeps Evolving
360 Zhihui Cloud Developer
360 Zhihui Cloud Developer
Sep 11, 2025 · Big Data

How Paimon Transforms Membership Data Warehousing: From Legacy Lambda to Real‑Time Lakehouse

This article examines the challenges of a legacy Lambda‑based membership data warehouse, introduces Apache Paimon’s lakehouse architecture and its key features, and showcases three real‑world implementations—partial‑update order wide tables, Bitmap‑based UV counting, and branch‑based data correction—while discussing benefits, remaining challenges, and future directions.

Data LakeFlinkPaimon
0 likes · 29 min read
How Paimon Transforms Membership Data Warehousing: From Legacy Lambda to Real‑Time Lakehouse
DataFunSummit
DataFunSummit
Sep 2, 2025 · Big Data

How Xiaomi Cuts Costs and Boosts Performance with Cloud‑Native Data Lake Architecture

Xiaomi’s engineers explain how they tackled data‑lake challenges—small files, metadata latency, and multi‑cloud costs—by combining compact storage, Gravitino‑based metadata governance, Iceberg and Paimon formats, and JuiceFS abstraction, achieving lower storage expenses, faster queries, and a roadmap toward intelligent, real‑time, multimodal lakehouses.

Data LakeMulti-Cloudbig data
0 likes · 14 min read
How Xiaomi Cuts Costs and Boosts Performance with Cloud‑Native Data Lake Architecture
Smart Sea Tide
Smart Sea Tide
Aug 25, 2025 · Big Data

Hudi vs Delta Lake vs Iceberg: Deep Technical Comparison of Data Lake Table Formats

This article provides a detailed technical comparison of Apache Hudi, Delta Lake, and Apache Iceberg, covering core features such as incremental pipelines, concurrency control, merge‑on‑read, partition evolution, multi‑mode indexing, ingestion tools, and real‑world use cases, and explains why Hudi often leads for heavy update workloads.

Apache HudiApache IcebergData Lake
0 likes · 14 min read
Hudi vs Delta Lake vs Iceberg: Deep Technical Comparison of Data Lake Table Formats
Big Data Technology Tribe
Big Data Technology Tribe
Aug 12, 2025 · Databases

Why Lakehouse Architecture Is Redefining Modern Data Platforms

This article explains the evolution from traditional data warehouses and data lakes to the unified Lakehouse architecture, detailing its design, benefits, challenges, and research directions for delivering high‑performance SQL and advanced analytics on open‑format storage.

Data LakeLakehouseMetadata Layer
0 likes · 20 min read
Why Lakehouse Architecture Is Redefining Modern Data Platforms
Past Memory Big Data
Past Memory Big Data
Jul 30, 2025 · Big Data

Why Iceberg Is Dropping Positional Deletes in Merge‑on‑Read Tables

The article explains how Apache Iceberg v3 replaces the scalable‑limited positional‑delete mechanism in Merge‑on‑Read tables with compact Deletion Vectors, detailing the performance, I/O and metadata drawbacks of positional deletes and showing how the new bitmap‑based approach resolves them.

Apache IcebergData LakeDeletion Vector
0 likes · 20 min read
Why Iceberg Is Dropping Positional Deletes in Merge‑on‑Read Tables
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jul 20, 2025 · Big Data

Exploring the Architecture of a Data Lake and Application Platform

This article outlines the overall architecture, data architecture, logical project structure, and the construction of a data resource center for a data lake and application platform, illustrated through a series of diagrams that depict each component and their interconnections.

Data LakeData Resource CenterSystem Design
0 likes · 1 min read
Exploring the Architecture of a Data Lake and Application Platform
DataFunSummit
DataFunSummit
Jul 18, 2025 · Big Data

Data Lake & Lakehouse Innovations: Real-Time Analytics and Industry Case Studies

This article presents a curated collection of cutting‑edge data lake and lakehouse case studies—including real‑time analytics, cloud‑native architectures, industry implementations from sales platforms to automotive IoT, and the latest advancements in open‑source projects—offering insights into modern big‑data strategies and governance.

Data LakeLakehouseReal-time Analytics
0 likes · 2 min read
Data Lake & Lakehouse Innovations: Real-Time Analytics and Industry Case Studies
DataFunSummit
DataFunSummit
Jul 12, 2025 · Big Data

How Fluss Unifies Stream and Lake to Power AI Data Pipelines

In the era of rapid AI growth, Fluss offers a unified lake‑stream architecture that tackles data quality, timeliness, scale, and multimodal challenges by tightly integrating Flink streaming with a high‑performance data lake, enabling seamless real‑time and batch analytics for AI workloads.

AIData LakeFlink
0 likes · 12 min read
How Fluss Unifies Stream and Lake to Power AI Data Pipelines
Big Data Technology & Architecture
Big Data Technology & Architecture
Jul 8, 2025 · Big Data

Flink’s AI Agents and Disaggregated State: Transforming Big Data

The article reviews key topics from the FFA2025 Singapore conference, highlighting Flink’s new AI‑focused Agents framework, the breakthrough Flink 2.0 disaggregated state architecture, emerging lake storage solutions like Paimon, and the Fluss streaming table store, illustrating how big‑data platforms are evolving for AI workloads.

AI AgentsData LakeDisaggregated State
0 likes · 6 min read
Flink’s AI Agents and Disaggregated State: Transforming Big Data
DataFunTalk
DataFunTalk
Jul 4, 2025 · Big Data

How Flink Agents and Flink 2.0 Are Powering Real‑Time AI at Scale

The Flink Forward Asia 2025 conference in Singapore showcased Apache Flink’s latest advances—including Flink Agents for system‑triggered AI, the cloud‑native Flink 2.0 with disaggregated state management, the multi‑modal lakehouse Paimon, and the Fluss table storage system—highlighting the ecosystem’s shift toward real‑time AI integration.

Apache FlinkData LakeFlink 2.0
0 likes · 9 min read
How Flink Agents and Flink 2.0 Are Powering Real‑Time AI at Scale
Baidu Geek Talk
Baidu Geek Talk
Jun 30, 2025 · Big Data

How Baidu’s Turing 3.0 Leverages Apache Iceberg to Boost Data Lake Performance

This article explains how Baidu’s next‑generation data platform Turing 3.0 integrates Apache Iceberg to solve the inefficiencies of the legacy MEG stack, detailing ecosystem components, migration strategies from Hive, table‑level optimizations, and future roadmap for high‑frequency, low‑latency analytics.

Apache IcebergData LakeHive Migration
0 likes · 17 min read
How Baidu’s Turing 3.0 Leverages Apache Iceberg to Boost Data Lake Performance
StarRocks
StarRocks
Jun 26, 2025 · Databases

What’s New in StarRocks 3.5? Snapshot Backup, Bulk Load, Partition & Transaction Enhancements

StarRocks 3.5 introduces a cluster‑level Snapshot backup for fast recovery, a bulk‑load optimization that reduces small files and compaction cost, smarter partition management with time‑based merging and TTL, multi‑statement transactions with full ACID guarantees, low‑cardinality dictionary support for lake tables, and several security and performance upgrades.

ACID TransactionsData LakeLow Cardinality Dictionary
0 likes · 17 min read
What’s New in StarRocks 3.5? Snapshot Backup, Bulk Load, Partition & Transaction Enhancements
DataFunSummit
DataFunSummit
Jun 18, 2025 · Big Data

How Real‑Time Lakehouse and Apache Paimon Transform Modern Data Architecture

This article explains the concept of a real‑time lakehouse, compares it with traditional batch warehouses, introduces Apache Paimon and its innovations such as native upserts, LSM storage, tags and branches, and showcases multiple enterprise use cases that demonstrate its low‑cost, low‑latency stream‑batch integration.

Apache PaimonData Lakereal-time lakehouse
0 likes · 17 min read
How Real‑Time Lakehouse and Apache Paimon Transform Modern Data Architecture
DataFunSummit
DataFunSummit
Jun 10, 2025 · Big Data

How OpenLake Redefines Data Lake Infrastructure for the AI Era

This article explores OpenLake's evolution as a data lake platform for AI, covering the transition from Hive to modern lake formats like Iceberg and Paimon, performance benchmarks, metadata management advances, intelligent storage optimization, and the integration of multimodal support with the Lance file format.

AIData LakeOpenLake
0 likes · 22 min read
How OpenLake Redefines Data Lake Infrastructure for the AI Era
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jun 10, 2025 · Big Data

Boosting Automotive Data Processing with Alibaba Cloud EMR Serverless Spark

This article details how a leading automotive parts supply‑chain platform migrated from a traditional Hadoop stack to Alibaba Cloud EMR Serverless Spark and DataWorks, achieving faster, more elastic, and cost‑effective data processing, enhanced AI integration, and significant operational improvements across multiple business scenarios.

Data LakeEMR ServerlessSpark
0 likes · 12 min read
Boosting Automotive Data Processing with Alibaba Cloud EMR Serverless Spark
DataFunTalk
DataFunTalk
Jun 4, 2025 · Artificial Intelligence

Coupang’s Distributed Cache Architecture Accelerates AI/ML Model Training

Coupang’s AI platform replaces costly data‑copy steps with a distributed cache that automatically pulls data from a central lake, boosts GPU utilization across regions, cuts storage and operational expenses, and speeds up model training by up to 40% while simplifying deployment via Kubernetes.

AIData LakeGPU
0 likes · 9 min read
Coupang’s Distributed Cache Architecture Accelerates AI/ML Model Training
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
May 19, 2025 · Industry Insights

How Xiaohongshu Built a Minute‑Level Near‑Real‑Time Data Warehouse with Incremental Computing

Facing billions of daily logs and the need for minute‑level experiment metrics, Xiaohongshu partnered with Yunqi Tech to design a generic incremental‑compute solution that delivers near‑real‑time data warehousing with lower cost, higher accuracy, simplified pipelines, and improved query performance.

Data LakeFlinkPaimon
0 likes · 24 min read
How Xiaohongshu Built a Minute‑Level Near‑Real‑Time Data Warehouse with Incremental Computing
Big Data Technology & Architecture
Big Data Technology & Architecture
May 16, 2025 · Big Data

Apache Gravitino: An Open‑Source Metadata Lake for Unified Data and AI Asset Management

Apache Gravitino is an open‑source metadata service platform that provides a unified, high‑performance, geographically distributed metadata lake, enabling end‑to‑end data governance, multi‑engine access, and direct management of both structured and unstructured data assets across diverse systems.

Apache GravitinoData Lakedata governance
0 likes · 9 min read
Apache Gravitino: An Open‑Source Metadata Lake for Unified Data and AI Asset Management
Tencent Cloud Developer
Tencent Cloud Developer
May 8, 2025 · Big Data

How Setats Unifies Stream, Batch, and Incremental Processing for Real‑Time Data Lakes

At the 2025 DA Data+AI Conference in Shanghai, Tencent Cloud unveiled Setats—a unified stream‑batch‑incremental engine that cuts system costs, delivers second‑level data visibility and real‑time changelog generation, and demonstrates measurable performance gains in automotive IoT analytics while integrating tightly with the WeData platform.

Batch ProcessingBig Data ArchitectureData Lake
0 likes · 5 min read
How Setats Unifies Stream, Batch, and Incremental Processing for Real‑Time Data Lakes
DataFunSummit
DataFunSummit
May 4, 2025 · Big Data

Iceberg Table Format Practice in Huawei Terminal Cloud

This article explains how Huawei's terminal cloud adopts the Apache Iceberg table format to efficiently manage large-scale datasets, detailing its architecture, feature engineering, merge operations, LSM-based storage, schema versioning, AB testing support, catalog enhancements, and future roadmap for full lifecycle data governance.

Data LakeHuawei CloudStreaming
0 likes · 13 min read
Iceberg Table Format Practice in Huawei Terminal Cloud
DataFunTalk
DataFunTalk
Apr 9, 2025 · Big Data

Highlights of the Apache Hudi Asia Technical Salon Hosted by Kuaishou – Practices and Innovations from Leading Companies

The Kuaishou‑hosted Apache Hudi Asia technical salon gathered over 230 attendees and featured seven experts from Kuaishou, Meituan, TikTok, Huawei, JD and others, who shared best practices, architecture designs, and performance optimizations for large‑scale data lake applications across AI, BI, and real‑time workloads.

AIApache HudiBatch Processing
0 likes · 14 min read
Highlights of the Apache Hudi Asia Technical Salon Hosted by Kuaishou – Practices and Innovations from Leading Companies
DataFunSummit
DataFunSummit
Apr 3, 2025 · Big Data

Apache Hudi Asia Technical Salon Highlights: Practices and Innovations from Kuaishou, Meituan, Douyin, Huawei, and JD

The Apache Hudi Asia technical salon held in Beijing on March 29 gathered over 230 on‑site participants and 16,000 online viewers, featuring expert talks from leading Chinese tech companies that showcased real‑world Hudi implementations, performance optimizations, and future roadmap for data‑lake technologies.

Apache HudiData LakeFlink
0 likes · 13 min read
Apache Hudi Asia Technical Salon Highlights: Practices and Innovations from Kuaishou, Meituan, Douyin, Huawei, and JD
Kuaishou Tech
Kuaishou Tech
Apr 2, 2025 · Big Data

Apache Hudi Asia Summit Successfully Held

The first Apache Hudi Asia Summit in Beijing attracted over 230 attendees, featuring technical discussions on data lake optimization and case studies from companies like Fastly and Meituan.

Apache HudiData LakeTechnical Conference
0 likes · 12 min read
Apache Hudi Asia Summit Successfully Held
AntData
AntData
Mar 20, 2025 · Big Data

Design and Optimization of Real‑time Data Lake Tables with Paimon and Flink for Advertising Diagnostics

This article presents a comprehensive exploration of using Apache Paimon and Flink to design lake tables that support minute‑level latency, low cost, and unified batch‑stream processing for advertising data, covering schema design, partitioning strategies, performance trade‑offs, cost analysis, and operational best practices.

Data LakeFlinkPaimon
0 likes · 34 min read
Design and Optimization of Real‑time Data Lake Tables with Paimon and Flink for Advertising Diagnostics
Alimama Tech
Alimama Tech
Mar 12, 2025 · Big Data

Design and Evolution of Alibaba Advertising Real-Time Data Warehouse

Alibaba Mama’s advertising platform migrated from a monolithic Flink‑Kafka pipeline to a layered Paimon lakehouse, adding DWS upsert support and multi‑layer storage, which delivers minute‑level data freshness, cuts latency by 2.5 hours, reduces resource use over 40 %, halves development effort and achieves ≥99.9 % availability.

AdvertisingAlibabaData Lake
0 likes · 18 min read
Design and Evolution of Alibaba Advertising Real-Time Data Warehouse
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Mar 6, 2025 · Big Data

Leveraging Apache Iceberg and AutoMQ for Real-Time Data Lake Ingestion: Architecture, Best Practices, and Cost Optimization

This article examines how Apache Iceberg’s snapshot‑based ACID transactions, logical‑physical partition evolution, and COW/MOR update modes enable efficient real‑time data lake ingestion, and demonstrates AutoMQ’s Kafka‑to‑Iceberg Table Topic solution that simplifies schema management, reduces latency, and cuts operational costs.

Apache IcebergAutoMQData Lake
0 likes · 14 min read
Leveraging Apache Iceberg and AutoMQ for Real-Time Data Lake Ingestion: Architecture, Best Practices, and Cost Optimization