Big Data

Showing 100 articles max
Data Integration and Governance
Data Integration and Governance
Jul 24, 2026 · Big Data

Data Mining Demystified: Principles, Process, and Methods Explained in One Guide

Enterprises often have abundant data but lack insight; this article explains how data mining moves beyond simple reporting to answer why events occur, predict future outcomes, and recommend actions, covering four analytical levels, core principles, a six‑step workflow, and common techniques with concrete business examples.

AnalyticsBusiness IntelligenceData Quality
0 likes · 18 min read
Data Mining Demystified: Principles, Process, and Methods Explained in One Guide
Smart Sea Tide
Smart Sea Tide
Jul 24, 2026 · Big Data

A Complete Data Analysis Knowledge Map

This article presents a comprehensive collection of visual knowledge maps that outline data analysis steps, foundational concepts, technical tools, business processes, analyst skill frameworks, and related topics such as Python learning paths, Excel formulas, MySQL, and statistical methods.

AnalyticsExcelPython
0 likes · 3 min read
A Complete Data Analysis Knowledge Map
Data Integration and Governance
Data Integration and Governance
Jul 22, 2026 · Big Data

Data Warehouse vs Data Mart vs Data Lake vs Data Middle Platform: Clear Differences

The article explains how data warehouses, data marts, data lakes, and data middle platforms each address distinct problems—unified analytics, departmental needs, raw data storage, and governed reusable capabilities—while outlining their relationships, typical use cases, implementation considerations, and guidance on choosing the right architecture for a given business stage.

Big DataData Martdata architecture
0 likes · 16 min read
Data Warehouse vs Data Mart vs Data Lake vs Data Middle Platform: Clear Differences
DataFunSummit
DataFunSummit
Jul 22, 2026 · Big Data

How Tencent Redefines Data Architecture for the Agent Era

With agents moving from Q&A to execution, traditional architectures expose three critical flaws—data stored in lakes, models in the cloud, and split scheduling—forcing petabyte‑scale data movement; Tencent Cloud’s big data AI DLC resolves this by running Spark and Ray side‑by‑side on the same lake, enabling closed‑loop processing and automatic trajectory capture.

AIAgentBig Data
0 likes · 2 min read
How Tencent Redefines Data Architecture for the Agent Era
ITPUB
ITPUB
Jul 21, 2026 · Big Data

Why Big Data Is Suddenly Falling Out of Favor

Although national data production reached 52.26 ZB in 2025 and continues to grow, the term “big data” is disappearing from strategic discussions because it no longer provides the organizational credit it once did, and enterprises now demand concrete value attribution, responsibility, and AI‑driven accountability.

AI impactBig DataData Governance
0 likes · 14 min read
Why Big Data Is Suddenly Falling Out of Favor
Subtle Storm
Subtle Storm
Jul 19, 2026 · Big Data

Core Big Data Technologies Explained with a Restaurant Analogy

The article breaks down the essential big data components—data ingestion, distributed storage, resource management, batch and stream processing, analytics, search, and cluster operations—using a restaurant metaphor and lists typical tools such as Flume, Kafka, HDFS, Spark, Flink, Elasticsearch, and Kubernetes.

Big DataCluster ManagementData Analytics
0 likes · 7 min read
Core Big Data Technologies Explained with a Restaurant Analogy
DataFunTalk
DataFunTalk
Jul 18, 2026 · Big Data

How Tencent Redefines Data Architecture for the Agent Era

The article analyzes how traditional data‑lake‑model‑cloud architectures expose three critical flaws for agentic AI—massive data movement, fragmented logging, and split compute—then details Tencent Cloud's Big Data AI DLC solution that unifies Spark and Ray on a single lake to enable in‑place processing, closed‑loop training, and cost‑effective iteration.

Agentic AIRaySpark
0 likes · 2 min read
How Tencent Redefines Data Architecture for the Agent Era
StarRocks
StarRocks
Jul 15, 2026 · Big Data

Building a Lake‑Stream Integrated Data Pipeline with Fluss, Paimon, and StarRocks

The article details how Taotian Group’s senior data engineer Zhu Ao designed a lake‑stream integrated architecture using Fluss for second‑level real‑time storage, Paimon for minute‑level lakehouse persistence, and StarRocks as a unified OLAP query layer, achieving over 50% faster development and more than 80% cost reduction.

Data IntegrationFlussLakehouse
0 likes · 14 min read
Building a Lake‑Stream Integrated Data Pipeline with Fluss, Paimon, and StarRocks
Data Integration and Governance
Data Integration and Governance
Jul 15, 2026 · Big Data

How to Build a Data Warehouse: End-to-End Process from Source Data to Analytics

Many companies start data‑warehouse projects by merely extracting data, building a few tables and adding a BI report, only to face inconsistent metrics, unreadable tables, mismatched numbers and endless Excel rechecks; the article outlines a full end‑to‑end process—from source‑data inventory and stable ingestion to layered storage, modeling, metric unification and quality monitoring—to ensure trustworthy, reusable analytics.

AnalyticsData ModelingData Quality
0 likes · 17 min read
How to Build a Data Warehouse: End-to-End Process from Source Data to Analytics
Data Integration and Governance
Data Integration and Governance
Jul 14, 2026 · Big Data

Why Every Data Asset Needs Catalog, Standards, Quality, and Lineage

Enterprises often build many systems and reports, but without a systematic data‑asset management framework—comprising a data catalog, unified standards, quality controls, and lineage tracing—issues like inconsistent metrics, unknown data sources, and hard‑to‑diagnose report errors persist, undermining reliable decision‑making.

Data Lineagedata asset managementdata catalog
0 likes · 12 min read
Why Every Data Asset Needs Catalog, Standards, Quality, and Lineage
DataFunTalk
DataFunTalk
Jul 14, 2026 · Big Data

How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based stack to a Lambda‑enhanced architecture and finally to a Lakehouse with incremental compute, cutting architecture complexity, resource and development costs by two‑thirds while delivering second‑level analytics on petabyte‑scale data.

Big DataClickHouseFlink
0 likes · 22 min read
How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era
Data Party THU
Data Party THU
Jul 12, 2026 · Big Data

Why Polars Beats Pandas for Massive Data Processing: A Deep Dive

Polars outperforms Pandas on large‑scale ETL by using a multithreaded lazy execution model, columnar Arrow storage, and query optimizations, delivering up to 94× speedups on 10 GB workloads, while Pandas remains suitable for smaller datasets and tight ML ecosystem integration.

Data EngineeringETLLazy Execution
0 likes · 15 min read
Why Polars Beats Pandas for Massive Data Processing: A Deep Dive
DataFunTalk
DataFunTalk
Jul 11, 2026 · Big Data

How Xiaohongshu’s Data Architecture Evolved for the Big AI Data Era

The article details Xiaohongshu’s journey from a simple ClickHouse‑based ad‑hoc analytics stack to a Lambda‑style architecture and finally to a lakehouse with generic incremental compute, cutting architecture complexity, resource cost and development effort each to roughly one‑third while achieving sub‑10‑second query latency on petabyte‑scale data.

AIBig DataClickHouse
0 likes · 21 min read
How Xiaohongshu’s Data Architecture Evolved for the Big AI Data Era
Data Party THU
Data Party THU
Jul 9, 2026 · Big Data

Big Data Challenge 2026: Monthly Star Winners Announced with Winning Teams’ Experience Shares (Third Edition)

The 2026 China University Computer Competition Big Data Challenge announced its Monthly Star winners, and the top teams detailed their data preparation, StockTransformer and LightGBM modeling pipelines, feature engineering, validation strategies, ensemble techniques, and key lessons learned from the competition.

Big Data CompetitionCross-Stock AttentionEnsemble Modeling
0 likes · 8 min read
Big Data Challenge 2026: Monthly Star Winners Announced with Winning Teams’ Experience Shares (Third Edition)
Data Integration and Governance
Data Integration and Governance
Jul 9, 2026 · Big Data

Data Warehouse vs Big Data Platform vs Data Lake vs Data Middle Platform vs Lake‑Warehouse Integration: What’s the Real Difference?

The article compares five data‑architecture concepts—data warehouse, big data platform, data lake, data middle platform, and lake‑warehouse integration—explaining the specific problems each solves, their core characteristics, advantages, risks, and guidance on when to adopt each solution.

Big Data PlatformData IntegrationLake‑Warehouse Integration
0 likes · 12 min read
Data Warehouse vs Big Data Platform vs Data Lake vs Data Middle Platform vs Lake‑Warehouse Integration: What’s the Real Difference?
Data Integration and Governance
Data Integration and Governance
Jul 8, 2026 · Big Data

How to Evaluate Data Asset Quality: Focus on Completeness, Accuracy, Consistency, and Timeliness

The article explains why data quality is critical for business value, defines the four core dimensions—completeness, accuracy, consistency, timeliness—details metrics and evaluation methods for each, presents case studies, outlines a weighted scoring model, and describes practical implementation steps and tool support for systematic data‑asset quality assessment.

Data AssetData GovernanceData Quality
0 likes · 19 min read
How to Evaluate Data Asset Quality: Focus on Completeness, Accuracy, Consistency, and Timeliness