Big Data

Showing 100 articles max
ByteDance Data Platform
ByteDance Data Platform
Aug 7, 2026 · Big Data

Can a Single AI‑Native Lakehouse Format Power Multimodal Data, Retrieval and Agent Memory?

At AICon 2026, Ma Jin explained how Lance’s layered lakehouse format—spanning file, table, index and catalog layers—addresses AI workloads by supporting wide columns, cheap schema evolution, mixed storage and low‑latency random reads, enabling unified handling of training data, RAG retrieval and Agent memory.

AI-native FormatAgent MemoryLakehouse
0 likes · 16 min read
Can a Single AI‑Native Lakehouse Format Power Multimodal Data, Retrieval and Agent Memory?
Cloud Architecture
Cloud Architecture
Aug 6, 2026 · Big Data

Exporting 10 Billion Elasticsearch Records: From Simple Script to Enterprise Offline Platform

The article analyses why exporting billions of Elasticsearch documents requires a full‑stack platform rather than a one‑off script, detailing the pitfalls of naive pagination, the benefits of PIT + search_after + slicing, and a complete architecture with Kafka, Redis, MySQL, Kubernetes and observability for reliable, scalable offline data export.

Big DataData ExportElasticSearch
0 likes · 40 min read
Exporting 10 Billion Elasticsearch Records: From Simple Script to Enterprise Offline Platform
Big Data Technology & Architecture
Big Data Technology & Architecture
Aug 5, 2026 · Big Data

12 Tough Data‑AI Interview Questions from Leading Companies – Can You Master the Latest Trends? (Part 1)

This article (part 1) presents twelve high‑level interview questions from top tech firms covering the shift from AI‑Ready to Data‑Agent‑Ready data warehouses, AI‑centric metadata and lineage, lakehouse formats like Paimon, and key Iceberg features that empower AI‑driven analytics.

AIData EngineeringIceberg
0 likes · 8 min read
12 Tough Data‑AI Interview Questions from Leading Companies – Can You Master the Latest Trends? (Part 1)
DataFunTalk
DataFunTalk
Aug 4, 2026 · Big Data

Recap of Tencent Cloud AI DLC Launch: Serverless Spark + Ray Integration for Agent‑Native Data Lake

The article reviews Tencent Cloud's AI DLC launch, detailing how the serverless Spark + Ray platform unifies data, compute, and agent workflows, introduces four architectural upgrades, showcases core engines (TCRay, Xpark, Meson, Open Engine), and presents benchmark results and real‑world practices from Bosch and WorkBuddy that demonstrate significant performance and productivity gains.

AI DLCBenchmarkRay
0 likes · 14 min read
Recap of Tencent Cloud AI DLC Launch: Serverless Spark + Ray Integration for Agent‑Native Data Lake
DataFunSummit
DataFunSummit
Aug 3, 2026 · Big Data

Why Real‑Time vs Batch Data Diverge 5% and Teams Revert to T+1: The Lambda Architecture Dilemma

Amid exploding real‑time data demand, the traditional Lambda architecture suffers from high cost, data inconsistency and operational complexity, prompting a shift to a unified incremental computation engine that delivers minute‑level latency, sub‑hourly cost, and sub‑1% result divergence, as demonstrated by Kuaishou and Xiaohongshu production deployments.

Big DataIncremental ComputationKuaishou
0 likes · 12 min read
Why Real‑Time vs Batch Data Diverge 5% and Teams Revert to T+1: The Lambda Architecture Dilemma
IT Services Circle
IT Services Circle
Aug 3, 2026 · Big Data

Is Pandas Still Viable in the Era of Massive Data?

As data volumes explode from megabytes to gigabytes and beyond, the article analyzes why traditional Pandas scripts falter, compares DuckDB, Polars, and Pandas on performance and memory usage, and proposes a layered pipeline that leverages each tool where it excels.

0 likes · 19 min read
Is Pandas Still Viable in the Era of Massive Data?
Xike
Xike
Aug 1, 2026 · Big Data

Big Data Series #6: Getting Started with Iceberg Lakehouse – Snapshots, Time Travel, and Writes

This tutorial explains how Apache Iceberg adds snapshot metadata to files on HDFS or object storage, enabling atomic writes, time‑travel queries, and schema evolution, and walks through a Docker‑based setup with a REST catalog, Spark 3.5.3, and hands‑on examples including table creation, data insertion, version queries, and troubleshooting tips.

Apache IcebergHDFSSchema Evolution
0 likes · 16 min read
Big Data Series #6: Getting Started with Iceberg Lakehouse – Snapshots, Time Travel, and Writes
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Aug 1, 2026 · Big Data

Designing a Differentiated DSP Near‑Real‑Time Data Warehouse with Flink, DLF Paimon, and EMR Serverless StarRocks

The DSP advertising data pipeline was re‑architected by splitting three distinct links—NRR, BT, and CT—and assigning DLF Paimon for massive low‑frequency data, EMR Serverless StarRocks primary tables for high‑value low‑latency queries, and unified service‑layer merging, achieving 2‑minute freshness, sub‑5 ms point queries, ~60% storage cost reduction, and improved fault isolation.

DLF PaimonDSPEMR Serverless
0 likes · 12 min read
Designing a Differentiated DSP Near‑Real‑Time Data Warehouse with Flink, DLF Paimon, and EMR Serverless StarRocks
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jul 31, 2026 · Big Data

Dual‑Dimension Cost Cutting for EMR Serverless Spark AI Functions

The article explains how EMR Serverless Spark AI Functions incur costs from model inference and Spark compute, and presents a two‑pronged cost‑saving strategy—AI query optimization to cut unnecessary calls and asynchronous Batch File inference to lower unit prices and release executor resources—complete with examples, benchmarks, and configuration guidance.

AI FunctionBatch InferenceEMR Serverless
0 likes · 20 min read
Dual‑Dimension Cost Cutting for EMR Serverless Spark AI Functions
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jul 30, 2026 · Big Data

How EMR Serverless Spark Achieves 4× Faster PB‑Scale Text Deduplication

The article analyzes how migrating a large‑scale text deduplication workflow to Alibaba Cloud EMR Serverless Spark, using built‑in MinHash‑LSH functions and the Fusion Engine vectorized executor, reduces processing time from days to hours, cuts shuffle failures to zero, and eliminates most operational overhead.

AI data preprocessingBig DataEMR Serverless Spark
0 likes · 16 min read
How EMR Serverless Spark Achieves 4× Faster PB‑Scale Text Deduplication
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jul 29, 2026 · Big Data

Exploring EMR Serverless StarRocks AI Functions: Multimodal Embedding, Semantic Aggregation, and Mixed Retrieval

The article analyzes the newly released AI Function suite in Alibaba Cloud EMR Serverless StarRocks, detailing multimodal embedding, AI‑driven aggregation, semantic filtering, mixed vector‑full‑text search, architectural advantages such as SQL‑native execution, async pipelines, bounded resources, and real‑world use cases in advertising, gaming, and finance.

AI FunctionSQLStarRocks
0 likes · 16 min read
Exploring EMR Serverless StarRocks AI Functions: Multimodal Embedding, Semantic Aggregation, and Mixed Retrieval
Java Architect Handbook
Java Architect Handbook
Jul 29, 2026 · Big Data

What Is Data Skew in Distributed Computing, Its Impact, and How to Fix It

The article defines data skew as an uneven key distribution after shuffle that makes a few tasks handle most of the data, explains the resulting slowdown, OOM, low resource utilization and job failure, and then details diagnosis methods and concrete solutions for group‑by, join, null‑value and framework‑level scenarios.

Data SkewDistributed ComputingHive
0 likes · 15 min read
What Is Data Skew in Distributed Computing, Its Impact, and How to Fix It
YiSu Grain
YiSu Grain
Jul 28, 2026 · Big Data

Data Warehouse vs Data Lake vs Lambda/Kappa: Choosing the Right Architecture

This article explains why online transaction databases (OLTP) should not be used for multi‑year analytics, outlines the four core characteristics of a data warehouse, details the ETL process, compares star and snowflake schemas, contrasts data warehouses with data lakes, and guides you through selecting Lambda or Kappa architectures using a concrete retail‑store scenario.

ETLKappa ArchitectureLambda Architecture
0 likes · 36 min read
Data Warehouse vs Data Lake vs Lambda/Kappa: Choosing the Right Architecture
YiSu Grain
YiSu Grain
Jul 28, 2026 · Big Data

Day 39: Big Data Architecture – Distributed Storage, Batch vs Stream Processing, and Compute‑Storage Separation

This lesson explains why a single database cannot scale for massive e‑commerce data, introduces the three core pillars of distributed storage—sharding, replication, and horizontal scaling—covers batch and stream processing differences with Spark, Hive, Flink and Storm, compares compute‑storage integration versus separation, and shows how Kafka, HDFS, Spark and Flink fit together in a real‑time data platform.

Big DataCompute‑Storage SeparationFlink
0 likes · 25 min read
Day 39: Big Data Architecture – Distributed Storage, Batch vs Stream Processing, and Compute‑Storage Separation
Smart Sea Tide
Smart Sea Tide
Jul 27, 2026 · Big Data

Comprehensive Overview of Big Data Open‑Source Frameworks

This article provides a detailed, structured survey of the most widely used open‑source big‑data technologies—including Hadoop ecosystems, storage systems, processing engines, query tools, data ingestion, exchange, messaging, scheduling, governance, visualization, mining, and cloud platforms—highlighting each project's origins, core features, typical use cases, and notable strengths or limitations to aid technology selection and system design.

Big DataFlinkHBase
0 likes · 59 min read
Comprehensive Overview of Big Data Open‑Source Frameworks