Tagged articles

Spark

648 articles · Page 1 of 7
Code Farming
Code Farming
Sep 4, 2026 · Big Data

How Big Data Platforms Turn Dormant Data into User Value: A Ride-Hailing Case Study

This article details a three-stage big data pipeline—collection, computation, and application—used by a ride-hailing platform to build user and vehicle profiles, enable intelligent order matching that boosts click-through rates by 30%, and create a self-reinforcing data flywheel that continuously improves service quality.

Data FlywheelData PipelineSpark
0 likes · 8 min read
How Big Data Platforms Turn Dormant Data into User Value: A Ride-Hailing Case Study
Big Data Technology Tribe
Big Data Technology Tribe
Aug 14, 2026 · Big Data

How lance‑spark Implements Blob V2 Support: A Deep Dive

The article explains how lance‑spark enables Lance's Blob V2 storage by marking a column with blob encoding and setting file_format_version ≥ 2.2, describes the metadata‑driven descriptor schema, the write path that still accepts Spark BINARY, and the size‑based placement of binary data into inline, packed, dedicated or external blob files.

ArrowBlob V2Data Lake
0 likes · 12 min read
How lance‑spark Implements Blob V2 Support: A Deep Dive
DataFunTalk
DataFunTalk
Aug 4, 2026 · Big Data

Recap of Tencent Cloud AI DLC Launch: Serverless Spark + Ray Integration for Agent‑Native Data Lake

The article reviews Tencent Cloud's AI DLC launch, detailing how the serverless Spark + Ray platform unifies data, compute, and agent workflows, introduces four architectural upgrades, showcases core engines (TCRay, Xpark, Meson, Open Engine), and presents benchmark results and real‑world practices from Bosch and WorkBuddy that demonstrate significant performance and productivity gains.

AI DLCBenchmarkData Lake
0 likes · 14 min read
Recap of Tencent Cloud AI DLC Launch: Serverless Spark + Ray Integration for Agent‑Native Data Lake
Xike
Xike
Aug 1, 2026 · Big Data

Big Data Series #6: Getting Started with Iceberg Lakehouse – Snapshots, Time Travel, and Writes

This tutorial explains how Apache Iceberg adds snapshot metadata to files on HDFS or object storage, enabling atomic writes, time‑travel queries, and schema evolution, and walks through a Docker‑based setup with a REST catalog, Spark 3.5.3, and hands‑on examples including table creation, data insertion, version queries, and troubleshooting tips.

Apache IcebergData LakeHDFS
0 likes · 16 min read
Big Data Series #6: Getting Started with Iceberg Lakehouse – Snapshots, Time Travel, and Writes
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jul 31, 2026 · Big Data

Dual‑Dimension Cost Cutting for EMR Serverless Spark AI Functions

The article explains how EMR Serverless Spark AI Functions incur costs from model inference and Spark compute, and presents a two‑pronged cost‑saving strategy—AI query optimization to cut unnecessary calls and asynchronous Batch File inference to lower unit prices and release executor resources—complete with examples, benchmarks, and configuration guidance.

AI FunctionEMR ServerlessSpark
0 likes · 20 min read
Dual‑Dimension Cost Cutting for EMR Serverless Spark AI Functions
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jul 30, 2026 · Big Data

How EMR Serverless Spark Achieves 4× Faster PB‑Scale Text Deduplication

The article analyzes how migrating a large‑scale text deduplication workflow to Alibaba Cloud EMR Serverless Spark, using built‑in MinHash‑LSH functions and the Fusion Engine vectorized executor, reduces processing time from days to hours, cuts shuffle failures to zero, and eliminates most operational overhead.

AI data preprocessingEMR Serverless SparkFusion Engine
0 likes · 16 min read
How EMR Serverless Spark Achieves 4× Faster PB‑Scale Text Deduplication
Java Architect Handbook
Java Architect Handbook
Jul 29, 2026 · Big Data

What Is Data Skew in Distributed Computing, Its Impact, and How to Fix It

The article defines data skew as an uneven key distribution after shuffle that makes a few tasks handle most of the data, explains the resulting slowdown, OOM, low resource utilization and job failure, and then details diagnosis methods and concrete solutions for group‑by, join, null‑value and framework‑level scenarios.

Data SkewDistributed ComputingHive
0 likes · 15 min read
What Is Data Skew in Distributed Computing, Its Impact, and How to Fix It
YiSu Grain
YiSu Grain
Jul 28, 2026 · Big Data

Day 39: Big Data Architecture – Distributed Storage, Batch vs Stream Processing, and Compute‑Storage Separation

This lesson explains why a single database cannot scale for massive e‑commerce data, introduces the three core pillars of distributed storage—sharding, replication, and horizontal scaling—covers batch and stream processing differences with Spark, Hive, Flink and Storm, compares compute‑storage integration versus separation, and shows how Kafka, HDFS, Spark and Flink fit together in a real‑time data platform.

Batch ProcessingCompute-Storage SeparationFlink
0 likes · 25 min read
Day 39: Big Data Architecture – Distributed Storage, Batch vs Stream Processing, and Compute‑Storage Separation
Smart Sea Tide
Smart Sea Tide
Jul 27, 2026 · Big Data

Comprehensive Overview of Big Data Open‑Source Frameworks

This article provides a detailed, structured survey of the most widely used open‑source big‑data technologies—including Hadoop ecosystems, storage systems, processing engines, query tools, data ingestion, exchange, messaging, scheduling, governance, visualization, mining, and cloud platforms—highlighting each project's origins, core features, typical use cases, and notable strengths or limitations to aid technology selection and system design.

FlinkHBaseHadoop
0 likes · 59 min read
Comprehensive Overview of Big Data Open‑Source Frameworks
DataFunSummit
DataFunSummit
Jul 22, 2026 · Big Data

How Tencent Redefines Data Architecture for the Agent Era

With agents moving from Q&A to execution, traditional architectures expose three critical flaws—data stored in lakes, models in the cloud, and split scheduling—forcing petabyte‑scale data movement; Tencent Cloud’s big data AI DLC resolves this by running Spark and Ray side‑by‑side on the same lake, enabling closed‑loop processing and automatic trajectory capture.

AIAgentData Lake
0 likes · 2 min read
How Tencent Redefines Data Architecture for the Agent Era
DataFunTalk
DataFunTalk
Jul 18, 2026 · Big Data

How Tencent Redefines Data Architecture for the Agent Era

The article analyzes how traditional data‑lake‑model‑cloud architectures expose three critical flaws for agentic AI—massive data movement, fragmented logging, and split compute—then details Tencent Cloud's Big Data AI DLC solution that unifies Spark and Ray on a single lake to enable in‑place processing, closed‑loop training, and cost‑effective iteration.

Agentic AIData LakeRay
0 likes · 2 min read
How Tencent Redefines Data Architecture for the Agent Era
DataFunTalk
DataFunTalk
Jul 14, 2026 · Big Data

How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based stack to a Lambda‑enhanced architecture and finally to a Lakehouse with incremental compute, cutting architecture complexity, resource and development costs by two‑thirds while delivering second‑level analytics on petabyte‑scale data.

ClickHouseFlinkIncremental Compute
0 likes · 22 min read
How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era
DataFunSummit
DataFunSummit
Jul 11, 2026 · Artificial Intelligence

Tencent CodeBuddy’s AI DLC Slashes Training Time and Costs with a Unified Spark‑Ray Service

The article explains how Tencent CodeBuddy’s AI DLC platform unifies Spark batch processing and Ray training to eliminate data movement, turning agent trajectories into reusable training fuel, which reduces monthly‑level training cycles to weekly, enables in‑place computation on billions of features, and cuts operational costs by 60%.

AI DLCData LakeGPU utilization
0 likes · 2 min read
Tencent CodeBuddy’s AI DLC Slashes Training Time and Costs with a Unified Spark‑Ray Service
DataFunTalk
DataFunTalk
Jul 6, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based ad‑hoc analysis system to a Lambda‑style architecture and finally to a lakehouse with incremental compute, cutting architecture complexity, resource and development costs by one‑third while delivering second‑level queries over petabyte‑scale data.

ClickHouseFlinkIncremental Compute
0 likes · 23 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
dbaplus Community
dbaplus Community
Jul 2, 2026 · Fundamentals

Production Hit by Silent Data Corruption: JDK 25 G1GC Bug Explained

A rare silent data‑corruption bug in JDK 25’s G1GC caused Parquet and ORC files written by Spark and Flink to become unreadable, prompting a multi‑stage investigation that traced the issue to an optional evacuation flaw affecting JNI‑pinned objects, which was later back‑ported and fixed in the OpenJDK community.

AI debuggingFlinkG1GC
0 likes · 20 min read
Production Hit by Silent Data Corruption: JDK 25 G1GC Bug Explained
DataFunTalk
DataFunTalk
Jun 30, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

Xiaohongshu, with over 3.5 billion monthly users and daily logs in the trillions, migrated 500 PB of data to Alibaba Cloud and iterated its data platform through four architecture generations—ClickHouse‑based ad‑hoc, Lambda, Lakehouse, and a unified incremental compute model—cutting resource, development, and storage costs to one‑third while delivering sub‑10‑second query latency at petabyte scale.

ClickHouseFlinkIncremental Compute
0 likes · 22 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
DataFunTalk
DataFunTalk
Jun 24, 2026 · Big Data

How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era

Xiaohongshu, with over 350 million monthly users and daily logs in the billions, migrated its data platform from AWS to Alibaba Cloud and iterated four times—from a ClickHouse‑based ad‑hoc layer to a Lambda architecture and finally a Lakehouse with incremental compute—cutting architecture complexity, resource cost and development effort each to about one‑third while delivering second‑level analytics on trillion‑scale data.

ClickHouseFlinkIncremental Compute
0 likes · 22 min read
How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era
Zhihu Tech Column
Zhihu Tech Column
Jun 22, 2026 · Big Data

How Zhihu Boosted Spark Job Performance with Gluten: A Practical Case Study

Zhihu migrated thousands of Spark SQL jobs to the open‑source Gluten/Velox engine, achieving up to 2.7× speedup, 44% CPU savings and 57% memory reduction while fixing numerous consistency, performance, and stability issues through custom optimizer rules and targeted patches.

Data ConsistencyGlutenPerformance Optimization
0 likes · 21 min read
How Zhihu Boosted Spark Job Performance with Gluten: A Practical Case Study
DataFunTalk
DataFunTalk
Jun 21, 2026 · Big Data

How Zhihu Optimized Spark Jobs with Gluten: A Practical Deep‑Dive

This article details Zhihu's end‑to‑end experience of migrating Spark SQL workloads to the open‑source Gluten framework, covering background performance benchmarks, the architecture of Gluten and Velox, consistency and performance challenges encountered during migration, the concrete fixes applied, and the resulting resource savings and future plans.

GlutenOptimizationPerformance
0 likes · 22 min read
How Zhihu Optimized Spark Jobs with Gluten: A Practical Deep‑Dive
DataFunTalk
DataFunTalk
Jun 20, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

The article details Xiaohongshu's step‑by‑step migration from a simple ClickHouse‑based analytics stack to a Lambda‑style 2.0 architecture and finally to a Lakehouse‑based 3.0 design, highlighting concrete performance numbers, cost reductions, and the definition of a generic incremental‑compute model (SPOT) that underpins the evolution.

ClickHouseFlinkIncremental Compute
0 likes · 22 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
DataFunTalk
DataFunTalk
May 28, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based ad‑hoc analysis to a Lambda‑style architecture and finally to a lakehouse with generic incremental compute, cutting architecture complexity, resource and development costs by one‑third while delivering second‑level queries over trillions of rows.

ClickHouseFlinkIncremental Compute
0 likes · 21 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
Big Data Tech Team
Big Data Tech Team
May 24, 2026 · Big Data

Data Warehouse Interview Pitfall Guide 2.0: Avoid Common SQL, Modeling, and ETL Mistakes

This guide compiles the most frequent interview pitfalls for data warehouse roles, covering SQL join and aggregation errors, window function misuse, subquery versus CTE performance myths, dimensional modeling mistakes, SCD implementation traps, layered design issues, data quality handling, ETL traps, Hive and Spark performance questions, real‑time warehousing considerations, and effective interview strategies.

ETLHiveInterview Preparation
0 likes · 3 min read
Data Warehouse Interview Pitfall Guide 2.0: Avoid Common SQL, Modeling, and ETL Mistakes
DataFunSummit
DataFunSummit
May 22, 2026 · Big Data

How OPPO Accelerates Multimodal Data & AI Fusion with Gravitino and Curvine

OPPO tackles explosive multimodal data growth by unifying metadata with Gravitino and boosting I/O performance using the open‑source Curvine cache, delivering a four‑layer data‑lake architecture that resolves data islands, metadata chaos, and bandwidth bottlenecks while achieving near‑commercial query speeds.

CurvineGravitinoLanceDB
0 likes · 11 min read
How OPPO Accelerates Multimodal Data & AI Fusion with Gravitino and Curvine
DataFunTalk
DataFunTalk
May 22, 2026 · Big Data

How Xiaohongshu Cut Data Architecture Complexity and Cost by One‑Third in the Big AI Data Era

The article details Xiaohongshu's evolution from a simple ClickHouse‑based analytics layer to a Lambda‑enabled 2.0 stack and finally a Lakehouse‑based 3.0 architecture, showing how each iteration reduced infrastructure complexity, resource consumption and development effort by roughly one‑third while supporting trillions of daily events and AI‑driven use cases.

ClickHouseFlinkIncremental Compute
0 likes · 21 min read
How Xiaohongshu Cut Data Architecture Complexity and Cost by One‑Third in the Big AI Data Era
DataFunTalk
DataFunTalk
May 11, 2026 · Big Data

How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based ad‑hoc analysis to a Lambda‑style architecture and finally to a lakehouse built on Iceberg, StarRocks, Flink and Spark, cutting architecture complexity, resource and development costs by two‑thirds while supporting trillions of daily events with sub‑second query latency.

ClickHouseFlinkIncremental Compute
0 likes · 22 min read
How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era
DataFunTalk
DataFunTalk
May 6, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

The article details Xiaohongshu's four‑stage data‑platform evolution—from a simple ClickHouse ad‑hoc setup to a Lambda‑based 2.0 design and finally a lakehouse‑driven 3.0 architecture—highlighting the adoption of general incremental compute, cost‑reduction to one‑third, performance gains of up to ten‑fold, and the SPOT standards that guide the new system.

ClickHouseFlinkIncremental Compute
0 likes · 21 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
DataFunTalk
DataFunTalk
Apr 29, 2026 · Big Data

How Xiaohongshu Revamped Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based analytics stack to a unified lakehouse with generic incremental compute, cutting architecture complexity, resource cost, and development effort by roughly one‑third while supporting petabyte‑scale, sub‑second queries across its 350 million‑user app.

ClickHouseFlinkIncremental Compute
0 likes · 22 min read
How Xiaohongshu Revamped Its Data Architecture for the Big AI Data Era
Lao Guo's Learning Space
Lao Guo's Learning Space
Apr 29, 2026 · Big Data

Designing a Full-Stack Credit Data System: From Ingestion to Real-Time Decision

The article dissects a credit data system architecture, detailing six logical layers—from multi-source data collection and feature engineering (including graph features and feature stores) to model training, real‑time stream processing, decision engine integration, and privacy‑preserving computation—while explaining the trade‑offs, tools, and performance targets needed for accurate, low‑latency risk assessment.

Credit ScoringData PipelineFeature Store
0 likes · 16 min read
Designing a Full-Stack Credit Data System: From Ingestion to Real-Time Decision
Big Data Tech Team
Big Data Tech Team
Apr 8, 2026 · Interview Experience

Master Spark Tuning for Data Warehouse Interviews: Real Cases & Tips

Learn how to demonstrate real Spark optimization skills in data‑warehouse interviews by exploring two detailed case studies—small‑file merging in ODS‑to‑DWD ETL and shuffle‑skew mitigation in DWS aggregation—plus key interview questions and practical troubleshooting steps that separate theory from hands‑on expertise.

Interview TipsSparkdata warehouse
0 likes · 9 min read
Master Spark Tuning for Data Warehouse Interviews: Real Cases & Tips
Baidu Geek Talk
Baidu Geek Talk
Mar 23, 2026 · Databases

How Baidu’s MEG Platform Revamped ClickHouse with a Lakehouse Architecture

This article analyzes the challenges of scaling ClickHouse within Baidu’s MEG data platform and details a lake‑house solution that decouples storage and compute, integrates a meta‑service for transparent data access, optimizes query performance through caching, data roll‑up and layout tuning, and introduces a unified query gateway that gracefully falls back to Spark for complex workloads.

ClickHouseLakehousePerformance Optimization
0 likes · 25 min read
How Baidu’s MEG Platform Revamped ClickHouse with a Lakehouse Architecture
Architect-Kip
Architect-Kip
Mar 2, 2026 · Big Data

How to Build a Scalable Tiered Archive & Query System for MySQL Data

This article presents a comprehensive design for a layered storage and unified scheduling platform that archives MySQL historical data, reduces storage costs, ensures high‑performance queries, and enables efficient data analysis through tiered hot, warm, and cold storage using big‑data technologies.

Data ArchivingDorisFlink
0 likes · 13 min read
How to Build a Scalable Tiered Archive & Query System for MySQL Data
Lakehouse Research Base
Lakehouse Research Base
Feb 12, 2026 · Big Data

Spark & Flink in Lakehouse: Latency, Accuracy & Cost Trade-offs for Batch-Stream Design

This article analyzes how Spark and Flink should divide responsibilities in lakehouse architectures by evaluating three core dimensions—data latency, computation accuracy, and resource cost—arguing that mature batch-stream integration balances business needs against technical overhead rather than pursuing universal real-time processing.

Batch ProcessingComputation AccuracyData Latency
0 likes · 9 min read
Spark & Flink in Lakehouse: Latency, Accuracy & Cost Trade-offs for Batch-Stream Design
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Feb 2, 2026 · Big Data

Real‑Time Analytics with Alibaba Cloud Serverless Spark & Paimon for Taobao Flash Sale

This article details how Alibaba Cloud EMR Serverless Spark combined with the Paimon lakehouse framework enables Taobao Flash Sale’s retail data team to achieve low‑latency, high‑throughput real‑time analytics, batch processing, and feature generation, outlining architecture evolution, performance gains, and practical Spark tuning techniques.

LakehousePaimonReal-time Analytics
0 likes · 18 min read
Real‑Time Analytics with Alibaba Cloud Serverless Spark & Paimon for Taobao Flash Sale
Lakehouse Research Base
Lakehouse Research Base
Jan 27, 2026 · Big Data

Big Data Expert's Failure in Traditional Enterprise: Heavy Tech Stack Wastes Resources on Small Data

A big data expert from a large tech company builds a full Hadoop/Spark/CDH stack for a traditional enterprise with only hundreds of thousands of daily records, causing high maintenance costs, half-hour query delays, and eventual departure; the case underscores the importance of matching technology to business scale, with StarRocks proposed as a lightweight alternative.

CDHHadoopSpark
0 likes · 11 min read
Big Data Expert's Failure in Traditional Enterprise: Heavy Tech Stack Wastes Resources on Small Data
Data Integration and Governance
Data Integration and Governance
Jan 20, 2026 · Big Data

Why Spark Isn't Enough: When to Use Other Tools in Your Data Projects

Although Spark excels at large‑scale, complex computations and offers high‑level APIs like DataFrame and SQL, real‑world projects often require additional tools for data ingestion, task scheduling, and simple streaming, so combining Spark with a data‑integration platform such as FineDataLink yields a more efficient and maintainable pipeline.

Data IntegrationData PipelineFineDataLink
0 likes · 9 min read
Why Spark Isn't Enough: When to Use Other Tools in Your Data Projects
Big Data Technology Tribe
Big Data Technology Tribe
Jan 20, 2026 · Big Data

Extending Spark SQL with LanceSparkSessionExtensions: A Complete Guide

This article explains how to inject the LanceSpark plugin into Spark, covering the core LanceSparkSessionExtensions class, various ways to register extensions, the custom parser and planner strategy implementations, and the underlying Spark mechanisms such as injectParser, injectPlannerStrategy, and PredicateHelper.

DataSourceV2LanceSparkPlannerStrategy
0 likes · 14 min read
Extending Spark SQL with LanceSparkSessionExtensions: A Complete Guide
ITPUB
ITPUB
Jan 15, 2026 · Databases

How to Migrate ClickHouse Data to Doris: Three Practical Strategies Tested

Facing a ClickHouse cluster shutdown, the author explores three migration methods—using Doris’s ClickHouse catalog, exporting to files with Broker/Stream Load, and Spark—to transfer ~10 billion rows to Doris, evaluating each for simplicity, bugs, and performance, and sharing detailed steps, code snippets, and benchmark results.

ClickHouseDorisSQL
0 likes · 9 min read
How to Migrate ClickHouse Data to Doris: Three Practical Strategies Tested
Data Integration and Governance
Data Integration and Governance
Jan 12, 2026 · Big Data

Choosing the Right Real‑Time Data Architecture: Tech Options and Performance Trade‑offs

The article explains that real‑time data processing requires not only speed but also accuracy, stability, and long‑term maintainability, compares Lambda, Kappa, and unified batch‑stream architectures, discusses lakehouse storage, and provides practical guidance on selecting the most suitable approach for different project needs.

FlinkKappa ArchitectureLakehouse
0 likes · 9 min read
Choosing the Right Real‑Time Data Architecture: Tech Options and Performance Trade‑offs
Big Data Tech Team
Big Data Tech Team
Jan 5, 2026 · Big Data

Top 10 Data Warehouse Interview Questions Every 2026 Engineer Must Master

This article compiles the most frequently asked interview questions for 2026 data‑warehouse development engineers, covering core concepts, layer architecture, SQL optimization, window functions, Hive vs Spark, data skew solutions, modeling metrics, slowly changing dimensions, scheduling tools, data quality monitoring, and real project experience.

Data ModelingHiveSQL optimization
0 likes · 8 min read
Top 10 Data Warehouse Interview Questions Every 2026 Engineer Must Master
DevOps Engineer
DevOps Engineer
Dec 27, 2025 · Artificial Intelligence

Demystifying GitHub AI: Models, Agents, Spaces, Spark, and More

This article explains GitHub's AI ecosystem—Models, Copilot, Agents, Spaces, Spark, Instructions, Skills, and the Model Context Protocol—clarifying each component, their relationships, and practical steps for developers to integrate them into their workflow.

AgentsCopilotGitHub AI
0 likes · 12 min read
Demystifying GitHub AI: Models, Agents, Spaces, Spark, and More
vivo Internet Technology
vivo Internet Technology
Dec 10, 2025 · Big Data

Vivo’s 800‑Day Journey Optimizing Celeborn Remote Shuffle Service at PB Scale

This technical report details how Vivo’s big‑data platform adopted Celeborn as its remote shuffle service, evaluated alternatives, tuned hardware and software configurations, implemented performance and stability enhancements, and outlines future operational and community‑driven improvements for handling petabyte‑scale shuffle workloads.

KubernetesRemote Shuffle ServiceSpark
0 likes · 20 min read
Vivo’s 800‑Day Journey Optimizing Celeborn Remote Shuffle Service at PB Scale
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Nov 15, 2025 · Big Data

From a Decade-Long Big Data Journey to a Cloud‑Native Lakehouse

This article chronicles a ten‑year evolution of a self‑built big data platform—detailing early Hadoop clusters, successive migrations to Spark, Hive, Hudi, and StarRocks, the operational challenges encountered, and the comprehensive shift to Alibaba Cloud EMR Serverless that delivered significant cost, performance, and stability gains while outlining future intelligent‑ecosystem plans.

Data LakeEMR ServerlessSpark
0 likes · 17 min read
From a Decade-Long Big Data Journey to a Cloud‑Native Lakehouse
Instant Consumer Technology Team
Instant Consumer Technology Team
Nov 10, 2025 · Big Data

Fixing Multi‑Version, Multi‑Cluster and HA with Apache Kyuubi for Spark/Flink

Apache Kyuubi, an enterprise‑grade multi‑tenant data gateway, replaces Livy and Flink SQL Gateway to support multiple engine versions, cross‑cluster elastic scheduling, high‑availability batch jobs, and traffic control, dramatically reducing deployment complexity, improving resource utilization, and accelerating release cycles for large‑scale Spark and Flink workloads.

Apache KyuubiData GatewayFlink
0 likes · 18 min read
Fixing Multi‑Version, Multi‑Cluster and HA with Apache Kyuubi for Spark/Flink
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Oct 18, 2025 · Big Data

Alibaba Cloud EMR’s AI Evolution: Accelerating Big Data Performance

Since its 2016 launch, Alibaba Cloud EMR has transformed from a basic open‑source Hadoop service into a high‑performance, AI‑enabled big‑data platform, delivering optimized I/O, vectorized processing, and integrated AI functions such as natural‑language SQL, StarRocks and Spark enhancements, while supporting diverse industry workloads.

Cloud ComputingEMRSpark
0 likes · 9 min read
Alibaba Cloud EMR’s AI Evolution: Accelerating Big Data Performance
Smart Sea Tide
Smart Sea Tide
Oct 9, 2025 · Big Data

Effective Spark Performance Tuning and Troubleshooting Guide

This article details practical Spark performance optimizations—including RDD reuse, broadcast variables, Kryo serialization, parallelism settings, shuffle tuning, and JVM tweaks—while also presenting systematic solutions for data skew, shuffle failures, serialization errors, and YARN mode issues, all illustrated with concrete code snippets and examples.

BroadcastCheckpointData Skew
0 likes · 25 min read
Effective Spark Performance Tuning and Troubleshooting Guide
Lakehouse Research Base
Lakehouse Research Base
Oct 3, 2025 · Operations

100 Advanced Linux Interview Questions for Big Data Operations

This comprehensive guide presents 100 advanced Linux interview questions tailored for big data operations, covering command-line mastery, system performance tuning, cluster configuration for Hadoop, Kafka, Spark, and Flink, complex troubleshooting scenarios, automation with Ansible and Prometheus, security hardening, and containerized deployments on Kubernetes.

HadoopKafkaLinux
0 likes · 45 min read
100 Advanced Linux Interview Questions for Big Data Operations
Lakehouse Research Base
Lakehouse Research Base
Oct 2, 2025 · Operations

100 Linux Interview Questions for Big Data Operations

This comprehensive guide presents 100 Linux interview questions tailored for big data operations, covering basic commands, permissions, process management, networking, log processing, shell scripting, and cluster-specific operations with detailed answers and practical explanations.

HadoopKafkaLinux
0 likes · 41 min read
100 Linux Interview Questions for Big Data Operations
DataFunSummit
DataFunSummit
Sep 21, 2025 · Big Data

Breaking the CPU Wall: BIGO’s Gluten Engine Accelerates Spark and Flink

When big‑data workloads hit the CPU wall, BIGO’s adoption of the open‑source Gluten project delivers native‑engine execution for Spark and a roadmap for Flink, achieving up to 30% end‑to‑end speedup, 50% memory savings, and a scalable, cost‑effective data processing platform.

FlinkGlutenSpark
0 likes · 16 min read
Breaking the CPU Wall: BIGO’s Gluten Engine Accelerates Spark and Flink
Architect's Must-Have
Architect's Must-Have
Sep 15, 2025 · Big Data

Mastering Spark Streaming Rate Control: A Deep Dive into Backpressure

This article explains Spark Streaming's rate control mechanisms, covering static limits, the dynamic back‑pressure feature introduced in Spark 1.5, the PID‑based estimator, RPC communication, and how Guava's token‑bucket RateLimiter enforces the calculated thresholds to ensure stability and optimal throughput.

BackpressureRateControlSpark
0 likes · 13 min read
Mastering Spark Streaming Rate Control: A Deep Dive into Backpressure
Big Data Tech Team
Big Data Tech Team
Aug 25, 2025 · Interview Experience

Essential Big Data Interview Questions for Data Warehouse Engineer Roles

A comprehensive list of interview topics covering self‑introduction, career moves, data‑warehouse design, team building, architecture comparisons, fact‑table classification, common dimensions, performance tuning, and data‑governance for aspiring big‑data engineers.

FlinkHiveInterview Preparation
0 likes · 4 min read
Essential Big Data Interview Questions for Data Warehouse Engineer Roles
Su San Talks Tech
Su San Talks Tech
Jul 17, 2025 · Big Data

How to De‑Duplicate 1 Billion QQ Numbers Using Under 1 GB of Memory

This article explores multiple techniques—including bitmap indexing, Bloom filters, external sorting, Spark, and layered bitmap structures—to efficiently remove duplicate QQ numbers from a dataset of up to one billion entries while keeping memory usage below a gigabyte and maintaining high accuracy.

BitmapBloom FilterSpark
0 likes · 12 min read
How to De‑Duplicate 1 Billion QQ Numbers Using Under 1 GB of Memory
Architect
Architect
Jul 7, 2025 · Big Data

How Baidu’s New Search Data Warehouse Architecture Boosts Performance by 5×

This article explains how Baidu’s search data team redesigned its data warehouse with wide‑table modeling, Parquet columnar storage, and a Spark‑ClickHouse fusion engine, eliminating redundancy, cutting query latency from minutes to seconds, and enabling self‑service analytics for thousands of users.

Data ModelingETLParquet
0 likes · 21 min read
How Baidu’s New Search Data Warehouse Architecture Boosts Performance by 5×
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jun 10, 2025 · Big Data

Boosting Automotive Data Processing with Alibaba Cloud EMR Serverless Spark

This article details how a leading automotive parts supply‑chain platform migrated from a traditional Hadoop stack to Alibaba Cloud EMR Serverless Spark and DataWorks, achieving faster, more elastic, and cost‑effective data processing, enhanced AI integration, and significant operational improvements across multiple business scenarios.

Data LakeEMR ServerlessSpark
0 likes · 12 min read
Boosting Automotive Data Processing with Alibaba Cloud EMR Serverless Spark
Lakehouse Research Base
Lakehouse Research Base
Apr 19, 2025 · Big Data

Essential Hadoop Ecosystem Commands: YARN, HDFS, HBase, Spark & Flink Reference

This comprehensive reference manual details essential command-line operations for core Hadoop ecosystem components including YARN resource management, HDFS distributed storage, HBase columnar database, Spark computing engine, and Flink stream processing framework, with practical examples for cluster administration, job submission, monitoring, and troubleshooting.

Cluster OperationsFlinkHBase
0 likes · 24 min read
Essential Hadoop Ecosystem Commands: YARN, HDFS, HBase, Spark & Flink Reference
Big Data Tech Team
Big Data Tech Team
Apr 17, 2025 · Big Data

Essential Spark Interview Q&A: Master Data Warehouse Engineer Questions

This article compiles a comprehensive set of Spark interview questions frequently asked by leading tech companies, providing detailed explanations of Spark’s performance mechanisms, architecture, RDD persistence, checkpointing, streaming, dependency types, HA setup, and practical coding examples to help data warehouse engineers prepare effectively.

RDDSparkSpark Streaming
0 likes · 21 min read
Essential Spark Interview Q&A: Master Data Warehouse Engineer Questions
vivo Internet Technology
vivo Internet Technology
Apr 16, 2025 · Big Data

Offline Mixed Deployment of Spark Tasks on Kubernetes: Containerization, Scheduling, and Elastic Resource Management

The article explains how the vivo Internet Big Data team containerized offline Spark jobs and deployed them with the Spark Operator on a mixed online‑offline Kubernetes cluster, using elastic scheduling and resource‑over‑subscription to boost CPU utilization by 30‑40% and handle over 100,000 daily tasks.

ContainerizationKubernetesSpark
0 likes · 36 min read
Offline Mixed Deployment of Spark Tasks on Kubernetes: Containerization, Scheduling, and Elastic Resource Management
DataFunSummit
DataFunSummit
Apr 3, 2025 · Big Data

Apache Hudi Asia Technical Salon Highlights: Practices and Innovations from Kuaishou, Meituan, Douyin, Huawei, and JD

The Apache Hudi Asia technical salon held in Beijing on March 29 gathered over 230 on‑site participants and 16,000 online viewers, featuring expert talks from leading Chinese tech companies that showcased real‑world Hudi implementations, performance optimizations, and future roadmap for data‑lake technologies.

Apache HudiData LakeFlink
0 likes · 13 min read
Apache Hudi Asia Technical Salon Highlights: Practices and Innovations from Kuaishou, Meituan, Douyin, Huawei, and JD
iQIYI Technical Product Team
iQIYI Technical Product Team
Mar 27, 2025 · Big Data

Cost‑Effective Real‑Time Data Warehouse 2.0: Migrating from Kafka to Iceberg

iQIYI transformed its real‑time data warehouse by replacing a costly Kafka‑based Lambda stack with a unified stream‑batch Iceberg lake, cutting storage expenses by 90%, halving compute costs, extending data retention, and delivering minute‑level freshness for 90% of use cases while preserving second‑level processing where needed.

FlinkKafkaReal-time Data Warehouse
0 likes · 11 min read
Cost‑Effective Real‑Time Data Warehouse 2.0: Migrating from Kafka to Iceberg
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Mar 20, 2025 · Big Data

How to Read and Write StarRocks Data with EMR Serverless Spark

This step‑by‑step guide explains how to use EMR Serverless Spark together with the StarRocks Spark Connector to create a workspace, upload the connector JAR, configure network connections, create databases and tables in StarRocks, and perform read/write operations via SQL sessions, Notebook sessions, or batch Spark jobs, complete with code examples and UI screenshots.

Data IntegrationEMR ServerlessSpark
0 likes · 14 min read
How to Read and Write StarRocks Data with EMR Serverless Spark
Big Data Technology & Architecture
Big Data Technology & Architecture
Mar 3, 2025 · Big Data

The Turning Point for Data Development: From Traditional Data Engineering to AI Data Engineering

The article analyzes how the rapid rise of open‑source large‑model AI in 2025 is reshaping the data development profession, urging developers to transition from specialized data‑engineer roles to full‑stack AI data engineering skills such as distributed computing, lake‑house architectures, and model tuning.

AIDistributed ComputingFlink
0 likes · 7 min read
The Turning Point for Data Development: From Traditional Data Engineering to AI Data Engineering
DataFunSummit
DataFunSummit
Feb 22, 2025 · Big Data

Blaze Engine: A Rust‑Based Native Vectorized Execution Engine for Spark SQL

The article introduces Blaze, Kuaishou's Rust‑powered native execution engine that vectorizes Spark SQL workloads, explains its architecture and operation, presents benchmark results showing up to 50% latency reduction, and details internal deployments, industry case studies, community collaborations, and the 2025 roadmap.

Performance OptimizationRustSpark
0 likes · 12 min read
Blaze Engine: A Rust‑Based Native Vectorized Execution Engine for Spark SQL
DataFunTalk
DataFunTalk
Feb 20, 2025 · Big Data

From Integrated Storage‑Compute to Decoupled Architecture: Practical Exploration of Kubernetes, Kyuubi, Celeborn, Blaze, and Hue in Big Data Platforms

This article analyzes the transition from a tightly coupled storage‑compute architecture to a decoupled model, detailing how Kubernetes, Kyuubi, Celeborn, Blaze, and Hue together solve resource inefficiencies, improve scalability, and boost query performance in modern big‑data environments.

BlazeKubernetesKyuubi
0 likes · 16 min read
From Integrated Storage‑Compute to Decoupled Architecture: Practical Exploration of Kubernetes, Kyuubi, Celeborn, Blaze, and Hue in Big Data Platforms
21CTO
21CTO
Feb 4, 2025 · Big Data

Why Python Beats Java and Scala for Modern Data Engineering

The article compares Java, Scala, SQL, and Python for data‑engineering tasks, arguing that Python’s versatility, rich ecosystem, and ease of use make it the preferred language for both small‑scale and massive Spark workloads despite its performance trade‑offs.

SQLScalaSpark
0 likes · 7 min read
Why Python Beats Java and Scala for Modern Data Engineering
DataFunSummit
DataFunSummit
Feb 1, 2025 · Big Data

Spark Native and Cloud Native: Vectorized SQL Engines, Remote Shuffle, and EMR Serverless Spark Practices

This article explains the challenges of big‑data processing in the cloud era, introduces Spark’s native‑language SQL engine rewrites, discusses vectorization and code generation techniques, describes cloud‑native storage‑compute separation with Remote Shuffle services such as Apache Celeborn, and presents the production benefits of Alibaba Cloud’s EMR Serverless Spark.

CodegenEMR ServerlessRemote Shuffle
0 likes · 12 min read
Spark Native and Cloud Native: Vectorized SQL Engines, Remote Shuffle, and EMR Serverless Spark Practices
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jan 26, 2025 · Big Data

How a FinTech Scaled Its Data Platform with Alibaba Cloud EMR Serverless Spark

Weifin, a fintech innovator, tackled massive data‑scale challenges by adopting Alibaba Cloud EMR Serverless Spark, building a unified Spark‑based platform that supports data collection, lake ingestion, distributed machine‑learning training, and intelligent risk‑control applications, while achieving performance gains, cost reduction, and scalable automation.

Sparkfintechmachine learning
0 likes · 10 min read
How a FinTech Scaled Its Data Platform with Alibaba Cloud EMR Serverless Spark
Airbnb Technology Team
Airbnb Technology Team
Jan 24, 2025 · Artificial Intelligence

Chronon — An Open-Source Framework for Production-Level Feature Engineering in Machine Learning

Chronon is an open‑source framework that centralizes feature definitions to guarantee training‑inference consistency, eliminates complex ETL pipelines, and supports real‑time and batch processing across diverse data sources, cutting feature‑development cycles from months to under a week, as demonstrated by Airbnb’s 40,000‑feature deployment.

ChrononHiveReal-time Data
0 likes · 10 min read
Chronon — An Open-Source Framework for Production-Level Feature Engineering in Machine Learning
dbaplus Community
dbaplus Community
Jan 19, 2025 · Big Data

How to Write Elegant, High‑Performance SQL for Big Data Pipelines

This article shares practical techniques for writing clean, efficient SQL in large‑scale data environments, covering predicate pushdown, sub‑queries, deduplication strategies, bucket optimization, and automation with Python‑Spark integration to improve readability and execution speed.

HiveOptimizationSpark
0 likes · 14 min read
How to Write Elegant, High‑Performance SQL for Big Data Pipelines
DataFunSummit
DataFunSummit
Jan 16, 2025 · Big Data

Zhihu Big Data Cost‑Reduction Practices: FinOps, Erasure Coding, ZSTD Compression, Spark Auto‑Tuning, and Remote Shuffle Service

This article details Zhihu's comprehensive cost‑reduction and efficiency‑boosting initiatives for its big‑data platform, covering FinOps‑driven financial operations, hybrid‑cloud architecture, cost allocation models, operational monitoring, and technical optimizations such as erasure coding, ZSTD compression, Spark auto‑tuning, and a remote shuffle service.

Cloud Cost ManagementFinOpsSpark
0 likes · 22 min read
Zhihu Big Data Cost‑Reduction Practices: FinOps, Erasure Coding, ZSTD Compression, Spark Auto‑Tuning, and Remote Shuffle Service
DataFunSummit
DataFunSummit
Jan 14, 2025 · Big Data

Tencent Real-Time Lakehouse Intelligent Optimization Practice

This presentation details Tencent's real‑time lakehouse architecture and the four key topics—lakehouse design, intelligent optimization services, scenario‑driven capabilities, and future outlook—covering components such as Spark, Flink, Iceberg, Auto‑Optimize Service, indexing, clustering, AutoEngine, and PyIceberg implementations.

Auto OptimizeFlinkLakehouse
0 likes · 12 min read
Tencent Real-Time Lakehouse Intelligent Optimization Practice
DataFunSummit
DataFunSummit
Jan 3, 2025 · Big Data

Tencent Real‑Time Lakehouse Intelligent Optimization Practices

This article presents Tencent's end‑to‑end real‑time lakehouse architecture, detailing its three‑layer design, the Auto Optimize Service modules such as compaction, indexing, clustering and engine acceleration, as well as scenario‑driven capabilities like multi‑stream joins, primary‑key tables, in‑place migration and PyIceberg support, and concludes with future optimization directions.

FlinkLakehouseSpark
0 likes · 11 min read
Tencent Real‑Time Lakehouse Intelligent Optimization Practices
Bilibili Tech
Bilibili Tech
Jan 3, 2025 · Big Data

Evolution and Production Practices of Apache Celeborn Remote Shuffle Service at Bilibili

Bilibili replaced Spark’s unstable External Shuffle Service with a push‑based approach, then deployed Apache Celeborn’s remote shuffle on Kubernetes using HA masters, tiered workers, extensive monitoring, history‑based routing, chaos testing, and seamless Spark, Flink, and MapReduce integration, while planning self‑healing, elastic scaling, and priority‑aware I/O enhancements.

Apache CelebornFlinkKubernetes
0 likes · 28 min read
Evolution and Production Practices of Apache Celeborn Remote Shuffle Service at Bilibili
Big Data Technology & Architecture
Big Data Technology & Architecture
Jan 2, 2025 · Big Data

Apache Paimon: Core Capabilities, Table Types, LSM Tree, Buckets, Merge Engines, and Operational Details

This article provides a comprehensive overview of Apache Paimon, covering its real‑time lake ingestion, unified stream‑batch processing, table types (primary‑key and append‑only), LSM‑tree storage, bucket mechanisms, merge‑engine options, compaction strategies, concurrency control, consumption methods, tag management, data cleanup, and system tables for big‑data workloads.

Apache PaimonFlinkLSM Tree
0 likes · 25 min read
Apache Paimon: Core Capabilities, Table Types, LSM Tree, Buckets, Merge Engines, and Operational Details
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 31, 2024 · Big Data

Eliminating Shuffle in Spark Joins with Storage Partitioned Join (SPJ) for Iceberg Tables

This article explains how Spark ≥ 3.3 introduces Storage Partitioned Join (SPJ) to avoid costly shuffle operations when joining partitioned V2 source tables such as Apache Iceberg, detailing the required conditions, configuration settings, practical code examples, and various join scenarios including mismatched partitions and data skew.

BucketingData SkewSQL
0 likes · 15 min read
Eliminating Shuffle in Spark Joins with Storage Partitioned Join (SPJ) for Iceberg Tables
JD Tech
JD Tech
Dec 30, 2024 · Big Data

Techniques for Writing Elegant and Efficient SQL in Big Data Environments

The article shares practical methods and code examples for making SQL both readable and high‑performing in large‑scale data platforms, covering predicate push‑down with subqueries, deduplication strategies, bucket utilization, and Python‑driven job parameter handling.

HivePerformanceSQL
0 likes · 14 min read
Techniques for Writing Elegant and Efficient SQL in Big Data Environments
DataFunSummit
DataFunSummit
Dec 27, 2024 · Big Data

Tencent Real-time Lakehouse Intelligent Optimization Practice

This presentation describes Tencent's real-time lakehouse architecture, including data lake compute, management, and storage layers, and details the intelligent optimization services—such as compaction, indexing, clustering, and auto-engine—designed to improve query performance, storage cost, and operational efficiency for large-scale data processing.

AutoEngineFlinkLakehouse
0 likes · 11 min read
Tencent Real-time Lakehouse Intelligent Optimization Practice
Bilibili Tech
Bilibili Tech
Dec 27, 2024 · Big Data

Consistency Architecture for Bilibili Recommendation Model Data Flow

The article outlines Bilibili’s revamped recommendation data‑flow architecture that eliminates timing and calculation inconsistencies by snapshotting online features, unifying feature computation in a single C++ library accessed via JNI, and orchestrating label‑join and sample extraction through near‑line Kafka/Flink pipelines, with further performance gains and Iceberg‑based future extensions.

Data ConsistencyFlinkProtobuf
0 likes · 12 min read
Consistency Architecture for Bilibili Recommendation Model Data Flow
Past Memory Big Data
Past Memory Big Data
Dec 27, 2024 · Big Data

How Uber Cuts Storage Costs with ZSTD Compression in Apache Parquet

Uber’s data lake on Hadoop stores hundreds of petabytes in Parquet files and, by adopting ZSTD compression, column pruning, and column reordering, achieves up to 79% storage reduction and significant vCore savings, with detailed benchmarks guiding optimal compression levels and open‑source contributions.

Apache ParquetCompressionHadoop
0 likes · 14 min read
How Uber Cuts Storage Costs with ZSTD Compression in Apache Parquet
Past Memory Big Data
Past Memory Big Data
Dec 26, 2024 · Big Data

Eliminate Shuffle: Deep Dive into Spark’s Storage Partition Join (SPJ)

This article explains how Spark ≥ 3.3’s Storage Partition Join (SPJ) can avoid costly shuffle operations by using Iceberg tables, outlines the required table properties and Spark configurations, demonstrates the effect with code examples and execution plans, and explores several realistic join scenarios.

Apache IcebergSPJShuffle Optimization
0 likes · 16 min read
Eliminate Shuffle: Deep Dive into Spark’s Storage Partition Join (SPJ)
dbaplus Community
dbaplus Community
Dec 24, 2024 · Big Data

How Bilibili Scaled Its Tag System for Massive Data and Real‑Time Accuracy

The article details Bilibili's comprehensive redesign of its tag system—including background challenges, architectural layers, technical upgrades like Iceberg integration and shard‑based ClickHouse writes, crowd selection methods, online service guarantees, performance metrics, and future plans—showcasing a data‑driven solution that boosts stability, speed, and business coverage.

ClickHouseDistributed ComputingOnline Service
0 likes · 24 min read
How Bilibili Scaled Its Tag System for Massive Data and Real‑Time Accuracy
Past Memory Big Data
Past Memory Big Data
Dec 24, 2024 · Big Data

Magnet: A Push‑Based Shuffle Service that Scales to Petabyte‑Level Data Processing

LinkedIn’s massive Spark workloads suffer from shuffle bottlenecks caused by tiny shuffle blocks, unreliable RPC connections, and data skew, so the authors design Magnet—a push‑merge shuffle service that merges blocks into large chunks, improves disk I/O, tolerates failures, and cuts end‑to‑end job time by nearly 30% regardless of hardware.

Disk I/O optimizationLarge‑scale data processingPush‑based service
0 likes · 56 min read
Magnet: A Push‑Based Shuffle Service that Scales to Petabyte‑Level Data Processing
dbaplus Community
dbaplus Community
Dec 14, 2024 · Databases

Why a Database‑First Operating System Could Replace Linux and Kubernetes

The article examines the DBOS concept—a database‑oriented operating system that places a distributed, transactional database at the core of OS services, tracing its roots from early database pioneers to modern cloud workloads and highlighting its potential advantages over traditional Linux‑Kubernetes stacks.

Cloud ComputingDBOSSpark
0 likes · 10 min read
Why a Database‑First Operating System Could Replace Linux and Kubernetes
Qunar Tech Salon
Qunar Tech Salon
Dec 10, 2024 · Big Data

Understanding and Solving Small File Problems in Hive and Spark

This article explains what constitutes a small file in HDFS, why they harm memory, compute and cluster load, outlines common sources such as data sources, streaming and dynamic partitioning, and provides detailed Hive and Spark solutions—including CombineHiveInputFormat, merge parameters, distribute by, and custom Spark extensions—to efficiently merge small files and improve job performance.

HiveMapReduceOptimization
0 likes · 23 min read
Understanding and Solving Small File Problems in Hive and Spark
Tongcheng Travel Technology Center
Tongcheng Travel Technology Center
Nov 27, 2024 · Big Data

Highlights of Tongcheng Travel’s 8th Big Data Technology Salon

The 8th Tongcheng Travel Big Data Technology Salon in Suzhou featured four expert talks covering Tencent Cloud’s Meson Spark engine, near‑line computing for travel itineraries, a Flink‑based real‑time risk control system, and Apache Paimon’s latest lake‑warehouse innovations, followed by a data‑driven business perspective session.

Apache PaimonData LakeFlink
0 likes · 7 min read
Highlights of Tongcheng Travel’s 8th Big Data Technology Salon
Bilibili Tech
Bilibili Tech
Nov 12, 2024 · Big Data

Scalable Tag System Architecture and Optimization

The rebuilt tag system introduces a three‑layer architecture, standard pipelines, Iceberg‑backed storage and custom ClickHouse sharding, a DSL for crowd selection, and a stateless online service, achieving 99.9% success, sub‑5 ms latency, and supporting thousands of tags across dozens of business scenarios while planning real‑time processing and automated lifecycle management.

ClickHouseOnline ServiceSpark
0 likes · 23 min read
Scalable Tag System Architecture and Optimization
Past Memory Big Data
Past Memory Big Data
Nov 8, 2024 · Big Data

How Spark on Kubernetes Transformed Duodian DMALL’s Big Data Platform

The article details Duodian DMALL’s migration from a traditional Hadoop stack to a cloud‑native Spark‑on‑Kubernetes architecture, explaining the motivations, design choices, component selections, operational challenges, and lessons learned through concrete examples and performance observations.

Apache CelebornFluent BitJuiceFS
0 likes · 21 min read
How Spark on Kubernetes Transformed Duodian DMALL’s Big Data Platform
Bilibili Tech
Bilibili Tech
Nov 1, 2024 · Big Data

Magnus: Intelligent Data Optimization Service for Iceberg Tables in Bilibili's Lakehouse Platform

Magnus is Bilibili’s self‑developed intelligent service that continuously optimizes Iceberg tables by scheduling snapshot expiration, orphan‑file cleanup, manifest rewriting, and multi‑dimensional data optimizations—including small‑file merging, sorting, distribution, and index creation—while automatically recommending configurations from real‑time query logs, delivering over 99.9% task success and up to 30% scan‑data reduction.

Data LakeIntelligent RecommendationOptimization
0 likes · 15 min read
Magnus: Intelligent Data Optimization Service for Iceberg Tables in Bilibili's Lakehouse Platform
Open Source Tech Hub
Open Source Tech Hub
Oct 31, 2024 · Big Data

How Bilibili Scaled Its Search Index with Distributed KV Storage and Spark

Bilibili transformed its search indexing pipeline by replacing a manual, low‑throughput process with a distributed KV store (Taishan) and Spark‑based construction, achieving unified data ingestion, reduced resource consumption, faster full‑ and incremental builds, and a shift from daily to hourly indexing cycles.

IndexingKV storeProtobuf
0 likes · 25 min read
How Bilibili Scaled Its Search Index with Distributed KV Storage and Spark