Tagged articles

big data

3804 articles · Page 20 of 39
DataFunSummit
DataFunSummit
Jan 23, 2022 · Big Data

MobTech's Integrated Data Governance Practices and Architecture

This article presents MobTech's comprehensive data governance and security practices, covering the necessity of governance, challenges in large‑scale data environments, the full‑link governance chain, modular architecture, and specific implementations for financial risk‑control scenarios.

big datadata architecturedata governance
0 likes · 19 min read
MobTech's Integrated Data Governance Practices and Architecture
DataFunTalk
DataFunTalk
Jan 22, 2022 · Big Data

Alibaba Cloud Data Integration (DataX) Architecture, Design Principles, and Solution Overview

This presentation details Alibaba Cloud DataWorks Data Integration (DataX), covering its architecture, core design principles, offline and real‑time synchronization mechanisms, deployment modes, product positioning, use‑case scenarios, and its role within the broader DataWorks ecosystem, highlighting its capabilities for large‑scale data movement and processing.

Alibaba CloudData IntegrationDataWorks
0 likes · 19 min read
Alibaba Cloud Data Integration (DataX) Architecture, Design Principles, and Solution Overview
Big Data Technology & Architecture
Big Data Technology & Architecture
Jan 18, 2022 · Big Data

Data Warehouse Data Quality Measurement Standards

The article outlines four key dimensions for evaluating data warehouse data quality—correctness, completeness, timeliness, and consistency—explains common consistency issues such as differing metric values across models, cross‑dimensional aggregations, and real‑time versus batch calculations, and proposes organizational and review mechanisms to mitigate these problems.

Data Qualitybig dataconsistency
0 likes · 9 min read
Data Warehouse Data Quality Measurement Standards
DataFunTalk
DataFunTalk
Jan 16, 2022 · Big Data

Time Series Database Capabilities and Application Scenarios in IoT, Smart Cities, and Edge Computing

This article explains the fundamentals of time‑series data, outlines the architecture and core technical advantages of Baidu Cloud's TSDB, and demonstrates how the database powers IoT, smart‑city, industrial, power‑grid, and autonomous‑driving use cases through multi‑level storage, distributed query optimization, and edge‑cloud integration.

Cloud ComputingData AnalyticsIoT
0 likes · 11 min read
Time Series Database Capabilities and Application Scenarios in IoT, Smart Cities, and Edge Computing
21CTO
21CTO
Jan 13, 2022 · Fundamentals

How to Achieve Data Maturity: Turning Data into a Strategic Product

The article explains why data maturity is essential for modern enterprises, defines its three pillars—people, tools, and readiness—shows how treating data as a product follows the same principles as great products, and outlines the four S (Speed, Scale, Simplicity, SQL) that guide a mature data ecosystem.

big datadata governancedata maturity
0 likes · 6 min read
How to Achieve Data Maturity: Turning Data into a Strategic Product
TAL Education Technology
TAL Education Technology
Jan 13, 2022 · Cloud Native

Offline Mixed Deployment with Kubernetes: Architecture, Implementation, and Performance Evaluation for Big Data Workloads

This article describes a cloud‑native offline mixed‑deployment solution that leverages Kubernetes to share resources between big‑data clusters and business services, outlines its implementation steps, presents detailed performance comparisons between Yarn and Kubernetes using TPC‑DS, Spark, and Terasort workloads, and discusses production experience and future plans.

Cloud NativeK8sKubernetes
0 likes · 8 min read
Offline Mixed Deployment with Kubernetes: Architecture, Implementation, and Performance Evaluation for Big Data Workloads
Shopee Tech Team
Shopee Tech Team
Jan 13, 2022 · Big Data

Engineering Practices and Performance Optimizations of Apache Druid for Real‑Time OLAP at Shopee

Shopee’s engineering team scaled a 100‑node Apache Druid cluster for real‑time OLAP by redesigning the Coordinator load‑balancing algorithm, adding incremental metadata pulls, introducing a segment‑merged result cache, and building exact‑count and flexible sliding‑window operators, while planning cloud‑native deployment.

Apache DruidBitmap IndexOLAP
0 likes · 17 min read
Engineering Practices and Performance Optimizations of Apache Druid for Real‑Time OLAP at Shopee
DataFunSummit
DataFunSummit
Jan 12, 2022 · Big Data

Exploring JD's Big Data Security and Distributed Permission System: Architecture, Principles, and Practices

This article presents JD's comprehensive big‑data security framework and distributed permission system, detailing the overall planning of the security center, data lifecycle protection strategies, core modules such as subjects, resources, policy language, and high‑performance access control, and how they address national compliance, business scalability, and technical challenges.

JD.comPermission Managementbig data
0 likes · 11 min read
Exploring JD's Big Data Security and Distributed Permission System: Architecture, Principles, and Practices
StarRocks
StarRocks
Jan 12, 2022 · Big Data

How Flink + StarRocks Deliver Lightning‑Fast Real‑Time Data Warehousing

This article explains the evolution, challenges, and technical solutions for building an end‑to‑end real‑time data warehouse by combining Apache Flink's stream processing with StarRocks' ultra‑fast OLAP engine, covering architecture, data models, integration methods, best‑practice cases, and future roadmap.

FlinkOLAPReal-time Data Warehouse
0 likes · 21 min read
How Flink + StarRocks Deliver Lightning‑Fast Real‑Time Data Warehousing
DataFunTalk
DataFunTalk
Jan 11, 2022 · Big Data

Interview with Wang Feng (Mo Wen): The Future of Apache Flink and Streaming Warehouses

In an exclusive InfoQ interview, Apache Flink community leader Wang Feng (aka Mo Wen) outlines the evolution of Flink toward a Streaming Warehouse, detailing recent technical advances, use‑case scenarios, and the upcoming Dynamic Table storage that aim to unify stream and batch processing for real‑time data‑warehouse workloads.

Apache FlinkDynamic TableFlink CDC
0 likes · 16 min read
Interview with Wang Feng (Mo Wen): The Future of Apache Flink and Streaming Warehouses
Big Data Technology & Architecture
Big Data Technology & Architecture
Jan 10, 2022 · Big Data

Key Takeaways from Flink Forward 2021: Real‑Time Computing, Flink SQL, ML, and Streaming Warehouse

The article reviews highlights from Flink Forward 2021, describing how real‑time computing is spreading across traditional industries, the unstoppable move toward Flink SQL, the emergence of Flink ML, and the vision of a streaming warehouse built on Flink Dynamic Table technology.

FlinkReal-Time ComputingStreaming Warehouse
0 likes · 8 min read
Key Takeaways from Flink Forward 2021: Real‑Time Computing, Flink SQL, ML, and Streaming Warehouse
Top Architect
Top Architect
Jan 9, 2022 · Information Security

Technical Analysis and Recent Updates of Xi'an “One Code Pass” System

The article reviews the Xi'an “One Code Pass” health‑code platform, covering its award recognition, recent service outages, capacity‑planning calculations, security‑platform procurement, Ministry engineer inspection, and the identified technical bottlenecks such as lack of CDN for static assets and insufficient outbound bandwidth.

One Code PassSystem ArchitectureXi'an
0 likes · 7 min read
Technical Analysis and Recent Updates of Xi'an “One Code Pass” System
21CTO
21CTO
Jan 8, 2022 · Big Data

How Amazon’s Intelligent Lakehouse Redefines Big Data Architecture

The article examines Amazon’s Intelligent Lakehouse architecture, tracing its evolution from early data‑lake‑warehouse integrations to a modern, serverless, secure, and AI‑enhanced platform that unifies data storage, governance, and analytics to lower big‑data costs and boost agility.

Data Lakebig datadata governance
0 likes · 12 min read
How Amazon’s Intelligent Lakehouse Redefines Big Data Architecture
DataFunTalk
DataFunTalk
Jan 8, 2022 · Big Data

Lakehouse: Concepts, Architecture, Implementation, and Cloud Practices

This article provides a comprehensive overview of the Lakehouse paradigm, tracing its origins from traditional data warehouses and data lakes, comparing architectures, detailing core components such as Delta Lake and Iceberg, and illustrating practical cloud implementations and future directions.

Apache IcebergCloud Data PlatformData Lake
0 likes · 14 min read
Lakehouse: Concepts, Architecture, Implementation, and Cloud Practices
Programmer DD
Programmer DD
Jan 8, 2022 · Big Data

How Flink’s Streaming Warehouse Is Redefining Real‑Time Data Lakes

This interview explores Apache Flink’s evolution toward a Streaming Warehouse, detailing its stream‑batch integration, new CDC‑based data integration, the Dynamic Table storage architecture, and how these innovations aim to simplify and accelerate real‑time big‑data analytics.

Apache FlinkDynamic TableFlink CDC
0 likes · 17 min read
How Flink’s Streaming Warehouse Is Redefining Real‑Time Data Lakes
HomeTech
HomeTech
Jan 6, 2022 · Operations

Design and Implementation of a Centralized Database Log Collection and Analysis Platform

This article describes the background, architecture, and implementation of a centralized database log collection and analysis platform built in 2021, detailing how logs from hosts, containers, and databases are normalized, streamed through Kafka, processed with Flink, stored in Elasticsearch, visualized with Kibana, and extended with alerting and configuration management to improve fault diagnosis and lay the groundwork for future AI‑driven operations.

KibanaLog Collectionbig data
0 likes · 5 min read
Design and Implementation of a Centralized Database Log Collection and Analysis Platform
Alibaba Cloud Developer
Alibaba Cloud Developer
Jan 6, 2022 · Big Data

Inside Alibaba Cloud’s MRACC Engine: How It Won the TPCx‑BB Benchmark

Alibaba Cloud’s self‑developed MRACC (Apasara Compute MapReduce Accelerator) leveraged hardware‑software integration, Spark and Hadoop optimizations, and eRDMA networking to achieve the top TPCx‑BB SF3000 performance, delivering up to 2‑3× faster SQL queries and 30% faster Spark shuffle, with significant cost efficiency gains.

BenchmarkRDMAbig data
0 likes · 9 min read
Inside Alibaba Cloud’s MRACC Engine: How It Won the TPCx‑BB Benchmark
Volcano Engine Developer Services
Volcano Engine Developer Services
Jan 4, 2022 · Big Data

How ByteDance Scales EB-Level Data: Architecture, BP Model & Real-Time Insights

ByteDance’s data platform, built over seven years, now handles exabyte-scale data and over 100 million TPS, using a hybrid “middle‑platform + Business Partner” model, custom engines like ClickHouse/ByteHouse, agile governance, and a suite of products to support internal and external businesses, illustrating large-scale big-data engineering practices.

ByteDanceClickHouseReal-time Analytics
0 likes · 22 min read
How ByteDance Scales EB-Level Data: Architecture, BP Model & Real-Time Insights
Big Data Technology & Architecture
Big Data Technology & Architecture
Jan 4, 2022 · Big Data

Big Data Mastery Roadmap: Learning Path, Resources, Future Trends and Interview Guidance

This comprehensive guide outlines a step‑by‑step learning roadmap for aspiring big data professionals, covering fundamentals, programming languages, Linux, databases, distributed theory, networking, offline and real‑time computing, data governance, warehouses, toolchains, video/book recommendations, future industry trends, interview tips, and community resources.

Interview TipsLearning PathReal-Time Computing
0 likes · 42 min read
Big Data Mastery Roadmap: Learning Path, Resources, Future Trends and Interview Guidance
DataFunTalk
DataFunTalk
Jan 3, 2022 · Databases

Pegasus: Architecture, New Features, Ecosystem, and Community Overview

This article introduces Pegasus, a distributed key‑value store, covering its background, system architecture, double‑WAL design, performance benchmarks, recent features such as hot backup, bulk load, access control, partition split, as well as its ecosystem tools and community development plans.

Hot BackupPEGASUSPartition Split
0 likes · 12 min read
Pegasus: Architecture, New Features, Ecosystem, and Community Overview
JavaEdge
JavaEdge
Jan 2, 2022 · Big Data

Mastering ZooKeeper: Core Concepts, Architecture, and Practical Setup

This article provides a comprehensive overview of ZooKeeper, covering its role in distributed systems, common use cases, source code setup, serialization and persistence mechanisms, network communication models, and the watcher workflow, enabling developers to understand and deploy ZooKeeper effectively.

PersistenceWatcherbig data
0 likes · 12 min read
Mastering ZooKeeper: Core Concepts, Architecture, and Practical Setup
DataFunTalk
DataFunTalk
Jan 1, 2022 · Big Data

JD's Flink Journey: Evolution, Optimizations, and Future Directions

This article details JD's adoption of Flink for real‑time computing, covering its evolution from Storm to Flink on Kubernetes, the platform architecture, major optimization techniques such as preview topology, backpressure handling, dynamic rebalance, checkpoint‑as‑savepoint, and outlines future plans including stream‑batch integration, stability improvements, intelligent operations, and AI integration.

FlinkJDKubernetes
0 likes · 10 min read
JD's Flink Journey: Evolution, Optimizations, and Future Directions
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 31, 2021 · Big Data

Apache SeaTunnel Joins the Apache Incubator: Overview, Features, and Real‑World Use Cases

SeaTunnel, the China‑originated data‑integration platform built on Spark and Flink, has been accepted into the Apache Incubator, and this article introduces its history, architecture, plugin ecosystem, deployment requirements, and numerous enterprise deployments across batch and streaming big‑data scenarios.

ApacheData IntegrationETL
0 likes · 7 min read
Apache SeaTunnel Joins the Apache Incubator: Overview, Features, and Real‑World Use Cases
IT Architects Alliance
IT Architects Alliance
Dec 31, 2021 · Industry Insights

A Complete 19‑Part Knowledge Map for Software Architects

The article presents a detailed 19‑section knowledge map for software architects, covering everything from core responsibilities and fundamentals to distributed caching, messaging, load balancing, performance testing, OS, algorithms, networking, databases, JVM, micro‑services, DDD, security, high availability, big data, and blockchain, with visual mind‑maps for each topic.

MicroservicesSystem Designbig data
0 likes · 4 min read
A Complete 19‑Part Knowledge Map for Software Architects
IT Architects Alliance
IT Architects Alliance
Dec 29, 2021 · Fundamentals

Collection of System Architecture Templates and Diagrams

This article presents a series of downloadable system architecture templates covering DMP, blockchain, data quality governance, enterprise technology, data architecture, Xelerator, alarm platform, microservices, front‑back separation, and a generic architecture, each illustrated with descriptive diagrams and brief explanations.

Architecture DiagramsMicroservicesSystem Architecture
0 likes · 5 min read
Collection of System Architecture Templates and Diagrams
Tencent Cloud Developer
Tencent Cloud Developer
Dec 28, 2021 · Industry Insights

How Flink and ClickHouse Combine to Build High‑Performance Real‑Time Data Warehouses

This article analyzes the challenges of massive data query efficiency, explains how Flink's stream processing and ClickHouse's OLAP engine complement each other, and presents a layered real‑time data‑warehouse architecture with practical guidance on data ingestion, write strategies, quality assurance, and evolving batch‑stream integration patterns.

ClickHouseFlinkOLAP
0 likes · 19 min read
How Flink and ClickHouse Combine to Build High‑Performance Real‑Time Data Warehouses
DataFunSummit
DataFunSummit
Dec 28, 2021 · Artificial Intelligence

Deep Application‑Driven Construction of Medical Knowledge Graphs: Methods, Models, and Case Studies

This article presents a comprehensive overview of medical knowledge graph development, covering global and domestic progress, domain characteristics, a six‑step construction workflow—including schema design, ontology term set creation, and graph building—and showcases practical applications such as intelligent alerts, guideline recommendations, and data direct reporting.

Data IntegrationMedical Knowledge Graphbig data
0 likes · 11 min read
Deep Application‑Driven Construction of Medical Knowledge Graphs: Methods, Models, and Case Studies
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 28, 2021 · Big Data

Comprehensive Guide to Spark SQL: Concepts, DataSet/DataFrame, Functions, Optimization and Common Pitfalls

This article provides an in‑depth overview of Spark SQL, covering its architecture, DataSet/DataFrame creation, DSL and SQL usage, integration with Hive, custom UDF/UDAF/Aggregator implementations, handling of small files, Cartesian product detection, and a catalog of useful built‑in functions and window operations.

DataFrameHiveSpark SQL
0 likes · 29 min read
Comprehensive Guide to Spark SQL: Concepts, DataSet/DataFrame, Functions, Optimization and Common Pitfalls
Su San Talks Tech
Su San Talks Tech
Dec 28, 2021 · Big Data

What Makes Kafka the Backbone of Real‑Time Big Data Processing?

This article provides a comprehensive overview of Apache Kafka, covering its distributed architecture, key advantages and drawbacks, the role of ZooKeeper, message delivery semantics, partitioning strategies, storage mechanisms, and performance optimizations such as zero‑copy and batch processing, all essential for high‑throughput real‑time data pipelines.

Distributed MessagingStreamingbig data
0 likes · 23 min read
What Makes Kafka the Backbone of Real‑Time Big Data Processing?
DataFunTalk
DataFunTalk
Dec 25, 2021 · Artificial Intelligence

Optimizing Spark‑ML Linear Models with Project Matrix: Background, Progress, and Future Plans

This article introduces the Project Matrix initiative that re‑examines and restructures Spark‑ML linear models, detailing the background of Spark‑ML usage at JD, the performance‑focused optimizations such as blockification and virtual centering, and outlines upcoming work to further improve scalability and accuracy.

Performance OptimizationSparkbig data
0 likes · 9 min read
Optimizing Spark‑ML Linear Models with Project Matrix: Background, Progress, and Future Plans
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 24, 2021 · Big Data

Key Updates and New Features in Apache Flink 1.14.2 Release

The Apache Flink 1.14.2 release, launched on December 16, fixes a critical Log4j vulnerability, resolves OOM issues with the Pulsar connector, introduces numerous Table API, DataStream API, connector, and checkpoint enhancements, deprecates several legacy APIs, and drops support for Apache Mesos, while also promoting related PDF resources.

Apache FlinkCheckpointsDataStream API
0 likes · 8 min read
Key Updates and New Features in Apache Flink 1.14.2 Release
AntTech
AntTech
Dec 23, 2021 · Databases

Understanding Graph Computing: Fundamentals, Applications, and Future Directions

This article explains graph computing fundamentals, illustrates its use in fraud detection, search ranking, and brain modeling, highlights Ant Group's record‑breaking performance and standards efforts, and outlines future challenges such as standardization, higher performance, and integration with AI.

Artificial IntelligenceGraph DatabasesPerformance
0 likes · 13 min read
Understanding Graph Computing: Fundamentals, Applications, and Future Directions
DataFunTalk
DataFunTalk
Dec 23, 2021 · Big Data

Building an Advertising Data Platform on ClickHouse: Architecture, Challenges, and Practices

This article details the design and implementation of an advertising data platform at eBay, explaining the business scenario, why ClickHouse was chosen over alternatives, the technical challenges faced, and the solutions involving lambda architecture, table engine choices, compression techniques, data ingestion pipelines, consistency guarantees, and deployment practices.

AdvertisingClickHouseData Ingestion
0 likes · 26 min read
Building an Advertising Data Platform on ClickHouse: Architecture, Challenges, and Practices
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 23, 2021 · Big Data

Key Spark Configuration Parameters and Their Explanations

This article presents a comprehensive list of essential Spark configuration settings—including executor memory, off‑heap memory, memory fractions, shuffle options, and adaptive query execution parameters—each accompanied by a concise description to help users fine‑tune Spark performance.

ShuffleSparkadaptive query execution
0 likes · 6 min read
Key Spark Configuration Parameters and Their Explanations
DataFunSummit
DataFunSummit
Dec 22, 2021 · Big Data

Data Governance Practices and Experiences at NetEase Cloud Music

This article details NetEase Cloud Music's comprehensive data governance journey, covering data warehouse architecture, data standards, event tracking (埋点) governance, asset lifecycle management, and future automation plans, illustrating how systematic governance improves data quality, cost efficiency, and business insight.

big datadata governancedata warehouse
0 likes · 21 min read
Data Governance Practices and Experiences at NetEase Cloud Music
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 22, 2021 · Big Data

Using Flink CDC to Capture MySQL Changes and Sink Them into ClickHouse

This article explains Change Data Capture (CDC), compares query‑based and log‑based approaches, introduces Debezium and ClickHouse, and provides step‑by‑step Flink CDC and Flink SQL CDC examples—including Java source, deserialization, sink code and required Maven dependencies—to stream MySQL binlog changes into ClickHouse for real‑time analytics.

CDCClickHouseData Streaming
0 likes · 14 min read
Using Flink CDC to Capture MySQL Changes and Sink Them into ClickHouse
Architects Research Society
Architects Research Society
Dec 21, 2021 · Fundamentals

Next-Generation Master Data Management (MDM): Architecture, Business Value, and Technical Challenges

This article explains master data management concepts, regulatory drivers, business benefits, key technical challenges, architectural trends such as graph databases and machine learning, and highlights leading vendors, providing a comprehensive overview for enterprises seeking modern MDM solutions.

AnalyticsMaster Data Managementbig data
0 likes · 9 min read
Next-Generation Master Data Management (MDM): Architecture, Business Value, and Technical Challenges
DataFunTalk
DataFunTalk
Dec 21, 2021 · Artificial Intelligence

Personalized Federated Learning and AI for Drug Discovery: Challenges, Applications, and Cloud Solutions

This talk by Huawei senior engineer Xu Chi explores the challenges of drug screening, AI-driven drug discovery practices, and how personalized federated learning combined with Huawei Cloud's high‑performance computing accelerates pharmaceutical research, including case studies, platform services, and collaborative efforts.

AICloud ComputingDrug discovery
0 likes · 11 min read
Personalized Federated Learning and AI for Drug Discovery: Challenges, Applications, and Cloud Solutions
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 21, 2021 · Big Data

Understanding Spark 3.0 Adaptive Query Execution (AQE) and Dynamic Partition Pruning (DPP)

This article explains the two most important Spark 3.0 features—Adaptive Query Execution and Dynamic Partition Pruning—detailing how AQE dynamically optimizes join strategies, partition coalescing, and skew handling, while DPP reduces I/O by pruning irrelevant fact‑table partitions at runtime.

Dynamic Partition PruningSQL optimizationSpark
0 likes · 10 min read
Understanding Spark 3.0 Adaptive Query Execution (AQE) and Dynamic Partition Pruning (DPP)
HelloTech
HelloTech
Dec 20, 2021 · Big Data

Building an ElasticSearch-based Search Platform for Ride-Hailing: Architecture, Data Synchronization, and Performance Optimization

Hello Mobility unified its fragmented ElasticSearch clusters into a single, real‑time search platform—leveraging Kafka‑driven CDC, Flink stream processing, custom ES plugins, and extensive performance tuning—to deliver scalable matching, recommendation and voice services, ultimately raising completed orders by 49.8 % and driver acceptance by 37 %.

FlinkSearch Platformbig data
0 likes · 19 min read
Building an ElasticSearch-based Search Platform for Ride-Hailing: Architecture, Data Synchronization, and Performance Optimization
Architecture Digest
Architecture Digest
Dec 20, 2021 · Backend Development

Understanding Kafka: Core Design, Architecture, and Performance

This article explains Kafka’s fundamental design concepts—including topics, partitions, replicas, consumer groups, and its network architecture—while highlighting performance features such as sequential writes, zero‑copy, log segmentation, and how the controller coordinates with ZooKeeper, providing a comprehensive overview for backend developers.

Backend DevelopmentKafkabig data
0 likes · 12 min read
Understanding Kafka: Core Design, Architecture, and Performance
DataFunSummit
DataFunSummit
Dec 18, 2021 · Big Data

Fast OLAP Forum – Latest Practices and Innovations in Real‑Time OLAP

The Fast OLAP Forum held on December 19 at DataFunCon gathers leading experts from Baidu, Tencent, JD, and FreeWheel to share cutting‑edge techniques in vectorized execution, cloud‑native ClickHouse, large‑scale OLAP architectures, and Presto optimizations, offering deep insights for practitioners dealing with massive real‑time data workloads.

Apache DorisClickHouseOLAP
0 likes · 7 min read
Fast OLAP Forum – Latest Practices and Innovations in Real‑Time OLAP
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 18, 2021 · Big Data

Slowly Changing Dimensions (SCD) – Design Principles, Challenges, and Hive Implementation

This article explains the concept of Slowly Changing Dimensions (SCD), discusses practical design questions, compares three change‑tracking requirements, presents three implementation patterns, and provides detailed Hive/SQL examples for historical data initialization and incremental updates in large‑scale data warehouses.

HiveSCDSQL
0 likes · 20 min read
Slowly Changing Dimensions (SCD) – Design Principles, Challenges, and Hive Implementation
Ctrip Technology
Ctrip Technology
Dec 16, 2021 · Big Data

Data Standard Management Practices in Ctrip Vacation Data Governance

This article outlines Ctrip Vacation's data standard management approach, covering why standards are needed, the three‑element framework of scope, tools, and policies, and detailed practices for data integration, production change handling, metadata governance, portal dashboard standardization, and self‑service query templating.

Data Integrationbig datadata governance
0 likes · 12 min read
Data Standard Management Practices in Ctrip Vacation Data Governance
High Availability Architecture
High Availability Architecture
Dec 16, 2021 · Big Data

iQIYI Basic Data Platform: Architecture, High Availability, and Service Practices

The iQIYI Basic Data Platform unifies internal data exchange standards, integrates massive multi‑business data, and implements high‑availability solutions for ID services, messaging, HBase storage, and read‑write scaling, showcasing practical engineering approaches to big‑data reliability and performance.

HBasebig datadistributed systems
0 likes · 11 min read
iQIYI Basic Data Platform: Architecture, High Availability, and Service Practices
政采云技术
政采云技术
Dec 16, 2021 · Big Data

What Is Event Tracking (埋点) and Its Implementation in a Data Analysis System

This article explains the concept of event tracking (埋点), its importance for capturing user behavior, outlines the four‑module architecture of a tracking system, compares code‑based, visual and full tracking methods, describes data models, storage, management, and presents a practical case study with analysis techniques.

AnalyticsBackendbig data
0 likes · 15 min read
What Is Event Tracking (埋点) and Its Implementation in a Data Analysis System
Liangxu Linux
Liangxu Linux
Dec 15, 2021 · Fundamentals

Cracking the 4‑Billion QQ Deduplication Challenge with 1 GB Memory

This article walks through four approaches—sorting, hashmap, file splitting, and a bitmap technique—to deduplicate 4 billion QQ numbers within a 1 GB memory limit, explains why the first three fail, and shows how a bitmap solves the problem efficiently.

BitmapMemory Optimizationalgorithm
0 likes · 8 min read
Cracking the 4‑Billion QQ Deduplication Challenge with 1 GB Memory
JD Cloud Developers
JD Cloud Developers
Dec 15, 2021 · Big Data

How JD Retail Scales Billion‑Item Selection with ClickHouse & Elasticsearch

This article details JD Retail's strategic "Nirvana" product‑selection platform, describing the technical challenges of handling billions of items and hundreds of tags, and presenting a dual‑engine solution using ClickHouse and Elasticsearch with Spark‑driven data pipelines to achieve fast filtering, multidimensional analytics, and efficient storage.

ClickHouseElasticsearchProduct Selection
0 likes · 15 min read
How JD Retail Scales Billion‑Item Selection with ClickHouse & Elasticsearch
DataFunSummit
DataFunSummit
Dec 14, 2021 · Big Data

Data Map: Background, Definition, and Youzan’s Practical Implementation

This article introduces the concept of a data map, explains its background and goals, describes Youzan’s end‑to‑end data‑map practice—including full data lineage, search, management, link analysis, impact estimation, and optimization—and concludes with a summary and future outlook.

big datadata governancedata lineage
0 likes · 16 min read
Data Map: Background, Definition, and Youzan’s Practical Implementation
Tencent Cloud Developer
Tencent Cloud Developer
Dec 13, 2021 · Cloud Computing

Trends in the Internet of Things and the Differentiated Development Path of Tencent Cloud IoT

Zhou Jiaxin explains that as IoT moves from universal connectivity to intelligent integration, Tencent Cloud IoT’s “Lianlian” platform tackles cost, efficiency and ecosystem gaps by embedding content, AI, big‑data and WeChat services into eight modular solutions, enabling rapid, cross‑industry smart applications.

AICloud ComputingIoT
0 likes · 13 min read
Trends in the Internet of Things and the Differentiated Development Path of Tencent Cloud IoT
Top Architect
Top Architect
Dec 13, 2021 · Big Data

Design and Implementation of BanYu's Big Data Access Control System

This article describes the evolution from an unsecured data warehouse to a comprehensive big‑data access control system at BanYu, detailing the background, data access methods, design goals, authentication and authorization mechanisms, policy configuration, integration with Metabase, and the overall workflow that balances security with efficiency.

Access ControlHiveLDAP
0 likes · 15 min read
Design and Implementation of BanYu's Big Data Access Control System
Python Crawling & Data Mining
Python Crawling & Data Mining
Dec 13, 2021 · Big Data

How to De‑duplicate 4 Billion QQ Numbers with Only 1 GB RAM

This article explains several algorithmic strategies—including sorting, hash maps, file splitting, and bitmap techniques—to remove duplicates from a file containing 4 billion QQ numbers while staying within a 1 GB memory limit, and it provides extension exercises for sorting, median, top‑K, and duplicate detection.

BitmapMemory Optimizationalgorithm
0 likes · 8 min read
How to De‑duplicate 4 Billion QQ Numbers with Only 1 GB RAM
Java Architect Essentials
Java Architect Essentials
Dec 11, 2021 · Information Security

Protecting Mobile Privacy in the Big Data Era: Risks of Data Leakage and How to Stay Safe

In today's big‑data era, excessive stress leads many to seek relief through risky online activities, but unauthorized app permissions and visits to dubious sites can expose personal information, so users must stay vigilant, limit permissions, avoid harmful sites, and use security tools to protect their mobile privacy.

big datadata leakagemobile security
0 likes · 6 min read
Protecting Mobile Privacy in the Big Data Era: Risks of Data Leakage and How to Stay Safe
IT Architects Alliance
IT Architects Alliance
Dec 11, 2021 · Big Data

Design and Implementation of Banyu's Big Data Permission System

This article describes the background, design goals, authentication and authorization mechanisms, system architecture, policy configuration, and Metabase integration of Banyu's big data permission system, which secures Hive, Presto, HDFS and other data access components using Apache Ranger and LDAP.

Access ControlApache RangerHive
0 likes · 14 min read
Design and Implementation of Banyu's Big Data Permission System
JD Retail Technology
JD Retail Technology
Dec 10, 2021 · Industry Insights

How JD Retail Cloud’s CRM Turned a Convenience Store Chain into a Data‑Driven Growth Engine

The award‑winning JD Retail Cloud store‑CRM solution helped the Haolinju convenience‑store chain overcome fragmented membership systems by rebuilding user data, applying big‑data algorithms and marketing automation, which boosted precise‑marketing ROI by 19% and increased purchase frequency by 0.7 per customer.

CRMCloud ComputingDigital Transformation
0 likes · 6 min read
How JD Retail Cloud’s CRM Turned a Convenience Store Chain into a Data‑Driven Growth Engine
DataFunTalk
DataFunTalk
Dec 10, 2021 · Big Data

Building and Evolving NetEase Yanxuan Real-Time Computing Platform: Architecture, SQLization, Serviceization, and Data Governance

This article details NetEase Yanxuan's real-time computing platform development from 2017 to present, covering its architecture, Flink‑SQL development environment, service‑oriented deployment, resource optimization, cloud‑native migration, comprehensive data governance, and future plans for stream‑batch integration and intelligent job diagnostics.

Cloud NativeFlinkReal-Time Computing
0 likes · 14 min read
Building and Evolving NetEase Yanxuan Real-Time Computing Platform: Architecture, SQLization, Serviceization, and Data Governance
DataFunSummit
DataFunSummit
Dec 10, 2021 · Big Data

Real‑Time Platform Construction at NetEase Yanxuan: Architecture, SQL‑Based Streaming, Serviceization, and Data Governance

This article details NetEase Yanxuan's evolution of a real‑time data platform from 2017 to present, covering background, current scale, layered architecture, Flink‑SQL development IDE, service‑oriented task execution, resource‑optimizing deployment modes, cloud‑native migration, comprehensive data governance, and future batch‑stream integration plans.

Cloud NativeFlinkbig data
0 likes · 15 min read
Real‑Time Platform Construction at NetEase Yanxuan: Architecture, SQL‑Based Streaming, Serviceization, and Data Governance
21CTO
21CTO
Dec 9, 2021 · Big Data

Designing a Scalable Big Data Permission System: From Hive to Metabase

BanYu’s early data warehouse lacked any access controls, prompting the creation of a comprehensive big‑data permission system that integrates authentication and authorization across Hive, Presto, HDFS, and Metabase using LDAP, Ranger policies, workflow automation, and both synchronous and asynchronous policy initialization.

AuthorizationHiveLDAP
0 likes · 16 min read
Designing a Scalable Big Data Permission System: From Hive to Metabase
DataFunTalk
DataFunTalk
Dec 9, 2021 · Big Data

Mobile Cloud LakeHouse: Cloud‑Native Big Data Analytics Architecture and Practices

This article introduces the cloud‑native LakeHouse solution from China Mobile Cloud, covering its lake‑warehouse integration concept, overall architecture, core functions such as storage‑compute separation, one‑click data ingestion, intelligent metadata discovery, serverless execution, JDBC support, incremental updates, and typical application scenarios in public and private clouds.

Cloud NativeData IntegrationKubernetes
0 likes · 17 min read
Mobile Cloud LakeHouse: Cloud‑Native Big Data Analytics Architecture and Practices
Architects' Tech Alliance
Architects' Tech Alliance
Dec 8, 2021 · Cloud Computing

Future Network Architecture and Emerging Data Center Technologies in the New Infrastructure Era

The article examines the concept of "new infrastructure" in China, outlines the evolution of the Internet toward a third generation, discusses candidate future network architectures such as SDN, CCN, and XIA, and reviews emerging data‑center networking technologies like large‑scale L2, VXLAN, virtual switching, and large‑scale switching, highlighting their role in supporting AI, big data, and cloud computing workloads.

AISDNbig data
0 likes · 12 min read
Future Network Architecture and Emerging Data Center Technologies in the New Infrastructure Era
Alimama Tech
Alimama Tech
Dec 8, 2021 · Big Data

Marketing Channel Attribution Models and Conversion Effectiveness Evaluation

Effective marketing budget allocation relies on robust channel attribution models that combine dimensions, metrics, and segmentation with rule‑based or data‑driven (Shapley) credit assignment across defined attribution windows, enabling multi‑touch analysis, conversion‑time insights, and ROI‑focused channel performance evaluation.

ROIattribution modelbig data
0 likes · 16 min read
Marketing Channel Attribution Models and Conversion Effectiveness Evaluation
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 8, 2021 · Big Data

Presto Overview, Architecture, and Query Optimization Techniques

This article introduces Presto, an open‑source MPP SQL engine, explains its coordinator‑worker architecture and connector model, and provides detailed storage, query, and join optimization strategies—including in‑memory parallelism, dynamic plan compilation, and practical SQL code examples—to achieve low‑latency, high‑performance analytics on big data.

Query OptimizationSQLbig data
0 likes · 7 min read
Presto Overview, Architecture, and Query Optimization Techniques
Open Source Linux
Open Source Linux
Dec 5, 2021 · Operations

Essential Skill Maps Every DevOps Engineer Should Master

This article compiles a series of visual skill maps covering DevOps, cloud computing, big data, security, architecture, and development practices, offering engineers a comprehensive roadmap to build and expand their technical knowledge across multiple domains.

Cloud ComputingDevOpsOperations
0 likes · 3 min read
Essential Skill Maps Every DevOps Engineer Should Master
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 4, 2021 · Big Data

Understanding Spark's BlockManager, MemoryStore, and DiskStore

This article explains Spark's storage architecture, detailing the roles and interactions of BlockManager, MemoryStore, and DiskStore, including their initialization, data management mechanisms, code implementations, and eviction strategies, to help readers grasp how Spark efficiently handles in‑memory and on‑disk data.

BlockManagerDiskStoreMemoryStore
0 likes · 12 min read
Understanding Spark's BlockManager, MemoryStore, and DiskStore
Open Source Linux
Open Source Linux
Dec 3, 2021 · Big Data

How Big Data Tech Evolved: Lessons from Alibaba, JD, and Didi

This article traces the evolution of big data technologies from early concepts and Google research papers through the rise of Hadoop, examines the platform transformations of Alibaba, JD.com, and Didi, and offers practical stack‑selection guidance for medium‑ and small‑scale enterprises.

AlibabaDidiJD.com
0 likes · 17 min read
How Big Data Tech Evolved: Lessons from Alibaba, JD, and Didi
21CTO
21CTO
Dec 2, 2021 · Fundamentals

Why China Is Betting on Open‑Source to Revitalize Its Software Industry

China's Ministry of Industry and Information Technology warns that the domestic software sector lags internationally and outlines a comprehensive plan—ranging from talent cultivation and open‑source community building to big‑data expansion—to transform the industry by 2025.

ChinaPolicybig data
0 likes · 5 min read
Why China Is Betting on Open‑Source to Revitalize Its Software Industry
Big Data Technology & Architecture
Big Data Technology & Architecture
Dec 1, 2021 · Big Data

Understanding Spark Shuffle: Mechanisms, Evolution, and Optimization

This article provides a comprehensive overview of Spark's shuffle process, explaining its definition, internal mechanisms such as shuffle write and read, the evolution of shuffle managers, and practical optimization techniques including parameter tuning and broadcast variables, all aimed at improving performance in large‑scale data processing.

ShuffleShuffle ReaderShuffle Writer
0 likes · 18 min read
Understanding Spark Shuffle: Mechanisms, Evolution, and Optimization
Alimama Tech
Alimama Tech
Dec 1, 2021 · Big Data

Optimization Algorithms for Guaranteed Delivery Advertising in Double‑11 Interactive Campaign

During Double‑11, the team created two specialized allocation algorithms—a brand‑score‑driven primal‑dual method for guaranteed‑downline contracts and a guarantee‑and‑balance flow‑re‑ranking approach for guaranteed‑non‑downline contracts—both using near‑line dual adjustments to meet contract volumes while boosting interaction depth, repeat visits, and browsing time.

Allocation AlgorithmOptimizationbig data
0 likes · 14 min read
Optimization Algorithms for Guaranteed Delivery Advertising in Double‑11 Interactive Campaign
Alibaba Cloud Developer
Alibaba Cloud Developer
Nov 30, 2021 · Big Data

Scaling Real‑Time Data Warehousing for Double‑11: Flink + Hologres in Action

During the 2021 Double‑11 shopping festival, logistics provider DiSiFang upgraded its real‑time data warehouse with Flink and Hologres, enabling multi‑billion‑row joins, cutting costs by 50%, and delivering stable, low‑latency analytics that powered high‑frequency dashboards and improved overall delivery speed.

Cloud ComputingFlinkHologres
0 likes · 13 min read
Scaling Real‑Time Data Warehousing for Double‑11: Flink + Hologres in Action
Qunar Tech Salon
Qunar Tech Salon
Nov 29, 2021 · Big Data

Construction and Practice of Qunar's Business Intelligence Platform

This article details the evolution, architecture, and technical choices of Qunar's BI platform—from early one‑stop reporting to a modular, self‑service system supporting real‑time analytics, multi‑metric calculations, and unified data governance—highlighting challenges, solutions, and performance benchmarks across big‑data technologies.

BIClickHouseReal-time Analytics
0 likes · 23 min read
Construction and Practice of Qunar's Business Intelligence Platform
Big Data Technology & Architecture
Big Data Technology & Architecture
Nov 28, 2021 · Big Data

OneData Methodology: Building a Unified Data Warehouse Architecture and Governance Framework

This article presents the OneData methodology for designing, standardizing, and governing a data warehouse, detailing background challenges, goals, industry references, core concepts, unified business and design consolidation, data modeling layers, naming conventions, data quality controls, and the resulting operational improvements and business value.

OneDatabig datadata governance
0 likes · 20 min read
OneData Methodology: Building a Unified Data Warehouse Architecture and Governance Framework
DataFunTalk
DataFunTalk
Nov 27, 2021 · Big Data

iQIYI Data Middle Platform: Architecture, Data Governance Practices, and Future Plans

The article details iQIYI’s data middle platform architecture and its comprehensive data governance practices, covering platform overview, data flow, unified standards, metadata management, production quality assurance, and future AI‑driven enhancements, illustrating how centralized data services improve reliability, efficiency, and security.

Data Qualitybig datadata governance
0 likes · 27 min read
iQIYI Data Middle Platform: Architecture, Data Governance Practices, and Future Plans
dbaplus Community
dbaplus Community
Nov 27, 2021 · Big Data

How Vipshop’s Hera Data Service Boosts Big Data Access and Performance

The article details the design, architecture, core features, scheduling logic, and performance gains of Vipshop’s self‑built Hera data service, which unifies data‑warehouse access, supports multiple engines, adapts SQL execution, and dramatically improves SLA for both B‑to‑B and B‑to‑C workloads.

Data ServiceDistributed ComputingETL
0 likes · 22 min read
How Vipshop’s Hera Data Service Boosts Big Data Access and Performance
Tencent Cloud Developer
Tencent Cloud Developer
Nov 26, 2021 · Big Data

WeChat's ClickHouse Real‑Time Data Warehouse: Challenges, Co‑Construction, and Performance Gains

Facing Hadoop’s minute‑to‑hour query latency on petabyte‑scale data, WeChat partnered with Tencent Cloud to build a ClickHouse‑based real‑time warehouse, adding custom ingestion, query‑optimisation and management tools that deliver billion‑row throughput, sub‑5‑second queries and over ten‑fold performance gains across millions of daily queries.

ClickHouseCloud NativeReal-time Analytics
0 likes · 9 min read
WeChat's ClickHouse Real‑Time Data Warehouse: Challenges, Co‑Construction, and Performance Gains
NiuNiu MaTe
NiuNiu MaTe
Nov 26, 2021 · Big Data

How to Deduplicate 4 Billion QQ Numbers Using Only 1 GB of Memory

This article walks through four practical techniques—sorting, hashmap, file splitting, and bitmap—to remove duplicate QQ numbers from a 4‑billion‑record file within a 1 GB memory limit, and provides extended exercises for sorting, median, top‑K, and duplicate detection.

Bitmapalgorithmbig data
0 likes · 8 min read
How to Deduplicate 4 Billion QQ Numbers Using Only 1 GB of Memory
StarRocks
StarRocks
Nov 24, 2021 · Big Data

Building a Scalable OLAP Platform at SF Express: StarRocks Evaluation and Lessons

SF Express’s data engineering team details how they migrated from a mixed‑component OLAP stack to a unified StarRocks platform, describing the evaluation criteria, performance‑critical design choices, import and query optimizations, and future roadmap for a high‑availability, low‑cost big‑data analytics solution.

OLAPSF ExpressStarRocks
0 likes · 14 min read
Building a Scalable OLAP Platform at SF Express: StarRocks Evaluation and Lessons
Qunar Tech Salon
Qunar Tech Salon
Nov 24, 2021 · Databases

Comprehensive Guide to Elasticsearch Index Design, Settings, and Mapping

This article provides a detailed guide on Elasticsearch index design, covering index settings, shard and replica planning, mapping strategies, complex types, lifecycle management, template usage, and practical best‑practice recommendations for large‑scale log data clusters.

ElasticsearchMappingSettings
0 likes · 27 min read
Comprehensive Guide to Elasticsearch Index Design, Settings, and Mapping
DataFunTalk
DataFunTalk
Nov 24, 2021 · Big Data

Tencent Game Big Data Analysis Engine: Architecture, Practices, and Future Plans

This article presents Tencent's game big‑data analysis platform, detailing its background, the architecture of the iData engine—including offline multi‑dimensional analysis (TGMars), online portrait analysis (TGFace), and real‑time multi‑dimensional analysis (TGDruid)—application scenarios, performance insights, and future ecosystem and open‑source plans.

Game AnalyticsOLAPReal-time Analytics
0 likes · 15 min read
Tencent Game Big Data Analysis Engine: Architecture, Practices, and Future Plans