Tagged articles

Big Data

3781 articles · Page 1 of 38
Smart Sea Tide
Smart Sea Tide
Aug 21, 2026 · Big Data

Designing Scalable Data Warehouses: Architecture, Modeling, Scheduling, and Metric Construction

The article explains how the surge of data in the DT era makes traditional storage insufficient, defines a data warehouse as a subject‑oriented, integrated, stable collection for decision support, outlines its development lifecycle—including integration, modeling, services, scheduling, metadata and quality management—and stresses that building a warehouse is an ongoing, iterative process driven by evolving business needs.

Big DataETLMetrics
0 likes · 3 min read
Designing Scalable Data Warehouses: Architecture, Modeling, Scheduling, and Metric Construction
Tencent Technical Engineering
Tencent Technical Engineering
Aug 20, 2026 · Big Data

Tencent SuperSQL Sets New TPC‑DS World Record, Leads in Performance and Cost Efficiency

Tencent's SuperSQL achieved a 654 million composite score on the TPC‑DS 100 TB benchmark—nearly ten times the previous record—while cutting performance‑per‑dollar cost to 11.04 CNY (about one‑sixth of earlier results), thanks to innovations in query optimization, a vectorized execution engine, adaptive memory and shuffle mechanisms, and AI‑driven diagnostics that together deliver superior speed, scalability, and business value.

AI DiagnosticsBig DataDistributed Computing
0 likes · 18 min read
Tencent SuperSQL Sets New TPC‑DS World Record, Leads in Performance and Cost Efficiency
Smart Sea Tide
Smart Sea Tide
Aug 18, 2026 · Big Data

How to Choose and Architect a Data Lake Platform for Enterprise Digital Transformation

The article outlines the strategic need for a unified data lake in a digital‑focused enterprise, details functional and non‑functional requirements such as linear scalability, real‑time and batch processing, multi‑tenant support, security and governance, and presents a comprehensive architecture design that integrates storage, compute, and management components.

Big DataData Lakecloud-native
0 likes · 19 min read
How to Choose and Architect a Data Lake Platform for Enterprise Digital Transformation
DataFunSummit
DataFunSummit
Aug 17, 2026 · Industry Insights

AI Era Data Infrastructure: From Storing Data to Enabling Agent‑Driven Context

The article analyzes how the rise of AI agents transforms data platforms from simple storage and query engines into AI‑native systems that provide trustworthy, real‑time context for autonomous decision‑making, outlining the three‑layer evolution of storage, compute, and application and the architectural upgrades required for modern data lakes.

AIAgentBig Data
0 likes · 13 min read
AI Era Data Infrastructure: From Storing Data to Enabling Agent‑Driven Context
Data Integration and Governance
Data Integration and Governance
Aug 11, 2026 · Big Data

What Is Data Architecture? Clarifying Databases, Data Warehouses, Lakes, and Middle Platforms

As enterprises add more systems, data volumes explode while accessing it becomes harder; this article defines data architecture, explains how databases, data warehouses, data lakes, and data middle platforms each solve distinct layers, and outlines the key questions and challenges for building a unified, reusable data ecosystem.

Big DataData LakeDatabase
0 likes · 13 min read
What Is Data Architecture? Clarifying Databases, Data Warehouses, Lakes, and Middle Platforms
AntData
AntData
Aug 7, 2026 · Big Data

Designing Lampara: Ant Group’s Real‑Time Data Processing System for End‑to‑End SLA Guarantees

The article details Ant Group’s Lampara, a next‑generation real‑time data processing platform that embeds end‑to‑end SLA guarantees, active disaster‑recovery, scenario‑driven development and enhanced operators, showing how it improves reliability, efficiency and cost while supporting large‑scale business scenarios such as flash sales, AI agents and marketing campaigns.

Big DataLamparaSLA
0 likes · 18 min read
Designing Lampara: Ant Group’s Real‑Time Data Processing System for End‑to‑End SLA Guarantees
Cloud Architecture
Cloud Architecture
Aug 6, 2026 · Big Data

Exporting 10 Billion Elasticsearch Records: From Simple Script to Enterprise Offline Platform

The article analyses why exporting billions of Elasticsearch documents requires a full‑stack platform rather than a one‑off script, detailing the pitfalls of naive pagination, the benefits of PIT + search_after + slicing, and a complete architecture with Kafka, Redis, MySQL, Kubernetes and observability for reliable, scalable offline data export.

Big DataData ExportElasticsearch
0 likes · 40 min read
Exporting 10 Billion Elasticsearch Records: From Simple Script to Enterprise Offline Platform
Smart Sea Tide
Smart Sea Tide
Aug 4, 2026 · Industry Insights

Why Are Fewer Companies Talking About Data Middle Platforms and Big Data Platforms Today?

The article explains that the decline of buzzwords like “big data platform” and “data middle platform” stems not from reduced data value but from a shift toward rational industry perception, highlighting the myth of scale, misaligned strategies, organizational challenges, and the amplified issues in the AI era.

AIBig Datadata governance
0 likes · 8 min read
Why Are Fewer Companies Talking About Data Middle Platforms and Big Data Platforms Today?
DataFunSummit
DataFunSummit
Aug 3, 2026 · Big Data

Why Real‑Time vs Batch Data Diverge 5% and Teams Revert to T+1: The Lambda Architecture Dilemma

Amid exploding real‑time data demand, the traditional Lambda architecture suffers from high cost, data inconsistency and operational complexity, prompting a shift to a unified incremental computation engine that delivers minute‑level latency, sub‑hourly cost, and sub‑1% result divergence, as demonstrated by Kuaishou and Xiaohongshu production deployments.

Big DataIncremental ComputationKuaishou
0 likes · 12 min read
Why Real‑Time vs Batch Data Diverge 5% and Teams Revert to T+1: The Lambda Architecture Dilemma
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jul 30, 2026 · Big Data

How EMR Serverless Spark Achieves 4× Faster PB‑Scale Text Deduplication

The article analyzes how migrating a large‑scale text deduplication workflow to Alibaba Cloud EMR Serverless Spark, using built‑in MinHash‑LSH functions and the Fusion Engine vectorized executor, reduces processing time from days to hours, cuts shuffle failures to zero, and eliminates most operational overhead.

AI data preprocessingBig DataEMR Serverless Spark
0 likes · 16 min read
How EMR Serverless Spark Achieves 4× Faster PB‑Scale Text Deduplication
YiSu Grain
YiSu Grain
Jul 29, 2026 · Fundamentals

45‑Minute Review: Master the Eight Architecture Types for the Soft Exam

This article guides readers through a 45‑minute review of the eight architecture categories—information‑system, layered, cloud‑native, SOA, embedded, communication, security, and big‑data—showing how to identify the relevant type from a problem statement, combine multiple architectures in a solution, and articulate the problem, solution, rationale, and cost with concrete examples and tables.

Big DataSOASoftware Architecture
0 likes · 32 min read
45‑Minute Review: Master the Eight Architecture Types for the Soft Exam
YiSu Grain
YiSu Grain
Jul 28, 2026 · Big Data

Day 39: Big Data Architecture – Distributed Storage, Batch vs Stream Processing, and Compute‑Storage Separation

This lesson explains why a single database cannot scale for massive e‑commerce data, introduces the three core pillars of distributed storage—sharding, replication, and horizontal scaling—covers batch and stream processing differences with Spark, Hive, Flink and Storm, compares compute‑storage integration versus separation, and shows how Kafka, HDFS, Spark and Flink fit together in a real‑time data platform.

Batch ProcessingBig DataCompute-Storage Separation
0 likes · 25 min read
Day 39: Big Data Architecture – Distributed Storage, Batch vs Stream Processing, and Compute‑Storage Separation
Smart Sea Tide
Smart Sea Tide
Jul 27, 2026 · Big Data

Comprehensive Overview of Big Data Open‑Source Frameworks

This article provides a detailed, structured survey of the most widely used open‑source big‑data technologies—including Hadoop ecosystems, storage systems, processing engines, query tools, data ingestion, exchange, messaging, scheduling, governance, visualization, mining, and cloud platforms—highlighting each project's origins, core features, typical use cases, and notable strengths or limitations to aid technology selection and system design.

Big DataFlinkHBase
0 likes · 59 min read
Comprehensive Overview of Big Data Open‑Source Frameworks
dbaplus Community
dbaplus Community
Jul 23, 2026 · Databases

Mid‑Year 2026 Database Roundup: AI‑Native Foundations and Major Product Updates

The article provides a detailed 2026 H1 review of global database trends, highlighting the shift to AI‑native and lake‑house architectures, summarizing major version releases across RDBMS, NoSQL, big‑data ecosystems and domestic solutions, and analyzing their technical features, compliance, performance and market implications.

AI-nativeBig DataCloud databases
0 likes · 37 min read
Mid‑Year 2026 Database Roundup: AI‑Native Foundations and Major Product Updates
Data Integration and Governance
Data Integration and Governance
Jul 22, 2026 · Big Data

Data Warehouse vs Data Mart vs Data Lake vs Data Middle Platform: Clear Differences

The article explains how data warehouses, data marts, data lakes, and data middle platforms each address distinct problems—unified analytics, departmental needs, raw data storage, and governed reusable capabilities—while outlining their relationships, typical use cases, implementation considerations, and guidance on choosing the right architecture for a given business stage.

Big DataData LakeData Mart
0 likes · 16 min read
Data Warehouse vs Data Mart vs Data Lake vs Data Middle Platform: Clear Differences
DataFunSummit
DataFunSummit
Jul 22, 2026 · Big Data

How Tencent Redefines Data Architecture for the Agent Era

With agents moving from Q&A to execution, traditional architectures expose three critical flaws—data stored in lakes, models in the cloud, and split scheduling—forcing petabyte‑scale data movement; Tencent Cloud’s big data AI DLC resolves this by running Spark and Ray side‑by‑side on the same lake, enabling closed‑loop processing and automatic trajectory capture.

AIAgentBig Data
0 likes · 2 min read
How Tencent Redefines Data Architecture for the Agent Era
ITPUB
ITPUB
Jul 21, 2026 · Big Data

Why Big Data Is Suddenly Falling Out of Favor

Although national data production reached 52.26 ZB in 2025 and continues to grow, the term “big data” is disappearing from strategic discussions because it no longer provides the organizational credit it once did, and enterprises now demand concrete value attribution, responsibility, and AI‑driven accountability.

AI impactBig Datadata governance
0 likes · 14 min read
Why Big Data Is Suddenly Falling Out of Favor
Subtle Storm
Subtle Storm
Jul 19, 2026 · Big Data

Core Big Data Technologies Explained with a Restaurant Analogy

The article breaks down the essential big data components—data ingestion, distributed storage, resource management, batch and stream processing, analytics, search, and cluster operations—using a restaurant metaphor and lists typical tools such as Flume, Kafka, HDFS, Spark, Flink, Elasticsearch, and Kubernetes.

Batch ProcessingBig DataCluster Management
0 likes · 7 min read
Core Big Data Technologies Explained with a Restaurant Analogy
DataFunTalk
DataFunTalk
Jul 14, 2026 · Big Data

How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based stack to a Lambda‑enhanced architecture and finally to a Lakehouse with incremental compute, cutting architecture complexity, resource and development costs by two‑thirds while delivering second‑level analytics on petabyte‑scale data.

Big DataClickHouseFlink
0 likes · 22 min read
How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era
DataFunTalk
DataFunTalk
Jul 11, 2026 · Big Data

How Xiaohongshu’s Data Architecture Evolved for the Big AI Data Era

The article details Xiaohongshu’s journey from a simple ClickHouse‑based ad‑hoc analytics stack to a Lambda‑style architecture and finally to a lakehouse with generic incremental compute, cutting architecture complexity, resource cost and development effort each to roughly one‑third while achieving sub‑10‑second query latency on petabyte‑scale data.

AIBig DataClickHouse
0 likes · 21 min read
How Xiaohongshu’s Data Architecture Evolved for the Big AI Data Era
DataFunTalk
DataFunTalk
Jul 6, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based ad‑hoc analysis system to a Lambda‑style architecture and finally to a lakehouse with incremental compute, cutting architecture complexity, resource and development costs by one‑third while delivering second‑level queries over petabyte‑scale data.

Big DataClickHouseFlink
0 likes · 23 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
DataFunSummit
DataFunSummit
Jul 4, 2026 · Industry Insights

How a Modern Data Platform Is Redefining the Future of Insurance

The article details how Ping An Property & Casualty transformed its legacy siloed data architecture into a systematic Kunpeng Intelligent Platform, built three core pillars—Agent platform, OSI semantic layer, and AI tools—boosted ChatBI accuracy, evaluated OpenClaw’s limits, and delivered end‑to‑end AI across marketing, underwriting, claims, agriculture, and forecasting.

AIBig DataEnd-to-End Automation
0 likes · 12 min read
How a Modern Data Platform Is Redefining the Future of Insurance
DataFunTalk
DataFunTalk
Jun 30, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

Xiaohongshu, with over 3.5 billion monthly users and daily logs in the trillions, migrated 500 PB of data to Alibaba Cloud and iterated its data platform through four architecture generations—ClickHouse‑based ad‑hoc, Lambda, Lakehouse, and a unified incremental compute model—cutting resource, development, and storage costs to one‑third while delivering sub‑10‑second query latency at petabyte scale.

Big DataClickHouseFlink
0 likes · 22 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jun 29, 2026 · Big Data

How DataWorks Data Agent Evolved Across Three Stages and Its Cloud‑Native Engineering Practices

The article systematically outlines DataWorks Data Agent’s progression from a Copilot‑assisted tool to human‑AI collaboration and finally AI‑driven autonomy, details its four‑agent product matrix covering data development, operations diagnostics, autonomous governance and ChatBI, describes three architecture iterations (Dify, AgentScope, QwenCode/OpenClaw) and a cloud‑managed deployment, and cites real‑world efficiency gains such as cutting development cycles from hours to minutes.

AI AgentBig DataData Agent
0 likes · 15 min read
How DataWorks Data Agent Evolved Across Three Stages and Its Cloud‑Native Engineering Practices
Smart Sea Tide
Smart Sea Tide
Jun 26, 2026 · Industry Insights

Ten Key Questions About Digital Twins

This article systematically examines ten fundamental questions on digital twins—defining the concept, identifying stakeholders, comparing global research intensity, linking to smart manufacturing, exploring integration with New IT, outlining scientific challenges, standards, and commercial tool requirements—to guide researchers, decision‑makers, and practitioners.

AIBig DataDigital Twin
0 likes · 23 min read
Ten Key Questions About Digital Twins
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jun 25, 2026 · Big Data

Taobao Live’s Shift from ETL to Managed Development with DataWorks Data Agent

The article details how Taobao Live’s data engineering team replaced traditional ETL bottlenecks with a three‑layer, AI‑native architecture built on DataWorks Data Agent, using NL2DSL2SQL, ontology‑driven knowledge bases, and multi‑agent collaboration to achieve near‑100% code generation and higher accuracy.

AI-nativeBig DataData Agent
0 likes · 10 min read
Taobao Live’s Shift from ETL to Managed Development with DataWorks Data Agent
DataFunTalk
DataFunTalk
Jun 25, 2026 · Big Data

From Writing SQL to Speaking Requirements: Practical Guide to DataWorks Data Agent

This article walks through using DataWorks Data Agent to automate end‑to‑end data‑warehouse development—from preparing source tables and a structured requirement document, uploading it, crafting task commands, selecting execution modes and models, to the agent generating SQL, building workflows, publishing them, and producing a final report—all without writing SQL manually.

AI automationBig DataData Agent
0 likes · 16 min read
From Writing SQL to Speaking Requirements: Practical Guide to DataWorks Data Agent
DataFunTalk
DataFunTalk
Jun 24, 2026 · Big Data

How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era

Xiaohongshu, with over 350 million monthly users and daily logs in the billions, migrated its data platform from AWS to Alibaba Cloud and iterated four times—from a ClickHouse‑based ad‑hoc layer to a Lambda architecture and finally a Lakehouse with incremental compute—cutting architecture complexity, resource cost and development effort each to about one‑third while delivering second‑level analytics on trillion‑scale data.

Big DataClickHouseFlink
0 likes · 22 min read
How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era
dbaplus Community
dbaplus Community
Jun 23, 2026 · Big Data

From Hand‑Written SQL to One‑Click Validation: Alibaba’s Verify‑Data Agent Skill Design Review

The article details how Alibaba’s production‑grade Verify‑Data Agent Skill replaces manual, multi‑SQL data validation with a single natural‑language command, automating table discovery, SQL generation, execution, and review‑level reporting, achieving up to 30‑minute turnaround, comprehensive coverage, and robust risk controls for big‑data pipelines.

Agent AutomationBig DataSQL Generation
0 likes · 28 min read
From Hand‑Written SQL to One‑Click Validation: Alibaba’s Verify‑Data Agent Skill Design Review
DataFunTalk
DataFunTalk
Jun 21, 2026 · Big Data

How Zhihu Optimized Spark Jobs with Gluten: A Practical Deep‑Dive

This article details Zhihu's end‑to‑end experience of migrating Spark SQL workloads to the open‑source Gluten framework, covering background performance benchmarks, the architecture of Gluten and Velox, consistency and performance challenges encountered during migration, the concrete fixes applied, and the resulting resource savings and future plans.

Big DataGlutenOptimization
0 likes · 22 min read
How Zhihu Optimized Spark Jobs with Gluten: A Practical Deep‑Dive
DataFunSummit
DataFunSummit
Jun 20, 2026 · Big Data

Building an Agentic Analytics Platform for the Gaming Industry with SelectDB

The article analyzes the fourfold challenges of game‑industry data analysis—high timeliness, massive concurrency, heterogeneous sources, and petabyte‑scale volumes—and explains how SelectDB’s evolution to an AI‑Ready, Agentic platform with MCP and a semantic layer addresses these issues through real‑time OLAP, multimodal processing, and autonomous decision loops.

AI-ReadyBig DataGame Data Analytics
0 likes · 16 min read
Building an Agentic Analytics Platform for the Gaming Industry with SelectDB
DataFunTalk
DataFunTalk
Jun 20, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

The article details Xiaohongshu's step‑by‑step migration from a simple ClickHouse‑based analytics stack to a Lambda‑style 2.0 architecture and finally to a Lakehouse‑based 3.0 design, highlighting concrete performance numbers, cost reductions, and the definition of a generic incremental‑compute model (SPOT) that underpins the evolution.

Big DataClickHouseFlink
0 likes · 22 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
DataFunSummit
DataFunSummit
Jun 19, 2026 · Big Data

Near‑Real‑Time Data Warehousing with Yunqi Lakehouse: Cases from Xiaohongshu, Kuaishou, Meituan

The article examines how Xiaohongshu, Kuaishou and Meituan adopted Yunqi Lakehouse’s General Incremental Computing and Single‑Engine architecture to achieve near‑real‑time data warehouses, cutting resource usage to as low as 1/20 of full‑batch jobs, reducing data latency from days to minutes, and improving query performance.

Big DataGeneral Incremental ComputingReal-time Data Warehouse
0 likes · 12 min read
Near‑Real‑Time Data Warehousing with Yunqi Lakehouse: Cases from Xiaohongshu, Kuaishou, Meituan
DataFunTalk
DataFunTalk
Jun 16, 2026 · Big Data

How MaxCompute Evolves Data Platforms for AI: Architecture, Features, and Real‑World Cases

The article explains how Alibaba Cloud's MaxCompute transforms a traditional data warehouse into a cloud‑native, multimodal Data+AI platform by introducing a four‑layer architecture, SQL‑based AI functions, the Python‑native MaxFrame framework, and a series of industry case studies that demonstrate performance gains and flexible resource scheduling.

Big DataData+AIMaxCompute
0 likes · 11 min read
How MaxCompute Evolves Data Platforms for AI: Architecture, Features, and Real‑World Cases
IT Learning Made Simple
IT Learning Made Simple
Jun 14, 2026 · Industry Insights

Why Data Architects Are the Hottest Talent in the DT Era

The article explains why data architects have become essential in the DT era, detailing their responsibilities, core skills, big‑data technology stack, governance practices, career paths, and the tools they use to turn data into a strategic asset for enterprises.

Big DataCareer Pathdata architecture
0 likes · 9 min read
Why Data Architects Are the Hottest Talent in the DT Era
dbaplus Community
dbaplus Community
Jun 14, 2026 · Big Data

Why Big Data Is Falling Silent: When Scale Can’t Fake Value Anymore

Although national data production reached 52.26 ZB in 2025 and keeps growing, the term “big data” is disappearing because it no longer serves as an organizational credit that hides the need for real value, responsibility, and measurable business impact, especially in the AI era.

AI impactBig Datadata governance
0 likes · 13 min read
Why Big Data Is Falling Silent: When Scale Can’t Fake Value Anymore
DataFunTalk
DataFunTalk
Jun 11, 2026 · Artificial Intelligence

How Qichacha Leverages Large Language Models for Field‑Level Data Lineage

This article details Qichacha's use of large language models to extract field‑level data lineage from heterogeneous, non‑standard code and ETL assets, describing the motivation, architectural blueprint, practical challenges such as cost, accuracy and hallucination, and the resulting improvements in impact analysis, metric tracing, and sensitive‑data governance.

Big DataFlinkLLM
0 likes · 11 min read
How Qichacha Leverages Large Language Models for Field‑Level Data Lineage
iQIYI Technical Product Team
iQIYI Technical Product Team
Jun 11, 2026 · Big Data

How iQIYI’s QBFS Enables Seamless Hybrid‑Cloud Storage and Cuts Big‑Data Costs by Over 30%

iQIYI’s big‑data team built a self‑developed QBFS virtual file system that unifies private and multiple public clouds, providing transparent routing, automatic migration, intelligent caching and fine‑grained governance, which together reduce storage and compute costs by more than 30 % while supporting scalable analytics.

Big DataData MigrationMulti-Cloud
0 likes · 21 min read
How iQIYI’s QBFS Enables Seamless Hybrid‑Cloud Storage and Cuts Big‑Data Costs by Over 30%
IT Learning Made Simple
IT Learning Made Simple
Jun 8, 2026 · R&D Management

The Essential Gear to Become a Software Architect

This guide maps the complete skill tree for aspiring software architects, detailing foundational knowledge, core competencies such as system design and performance tuning, extended expertise in cloud‑native and big‑data technologies, and a staged learning roadmap to help newcomers acquire the necessary gear.

Big DataSoftware ArchitectureSystem Design
0 likes · 9 min read
The Essential Gear to Become a Software Architect
DataFunSummit
DataFunSummit
Jun 7, 2026 · Artificial Intelligence

How Qichacha Uses Large Language Models for Field‑Level Data Lineage

This article details Qichacha's technical journey of applying large language models to resolve field‑level data lineage challenges in a complex, multi‑source data environment, describing the motivation, architecture, practical implementation, engineering trade‑offs, and measurable outcomes.

AIBig DataFlink
0 likes · 11 min read
How Qichacha Uses Large Language Models for Field‑Level Data Lineage
Digital Planet
Digital Planet
Jun 6, 2026 · Big Data

Why Has the Term “Big Data” Suddenly Disappeared?

Although data production continues to surge—reaching 52.26 ZB in 2025—the “big data” label is fading because its original narrative of scale as value has run out, exposing a credit‑and‑responsibility gap that forces organizations to demand concrete business impact rather than mere infrastructure.

AI impactBig Datadata governance
0 likes · 15 min read
Why Has the Term “Big Data” Suddenly Disappeared?
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
Jun 4, 2026 · Big Data

Scalar‑Vector Hybrid Search in a Data Lake with One SQL on EMR Serverless Spark

EMR Serverless Spark now supports scalar‑vector hybrid search via DLF Global Index, allowing a single Spark SQL statement to perform vector similarity and scalar filtering together, eliminating data movement, reducing latency, and boosting performance for scenarios such as autonomous driving, e‑commerce, and knowledge‑base retrieval.

Big DataDLF Global IndexEMR Serverless Spark
0 likes · 17 min read
Scalar‑Vector Hybrid Search in a Data Lake with One SQL on EMR Serverless Spark
Smart Sea Tide
Smart Sea Tide
Jun 2, 2026 · Big Data

Common Big Data Collection Tools and Their Core Features

This article reviews seven widely used big data collection tools—Flume, Fluentd, Logstash, Chukwa, Scribe, Splunk, and Scrapy—detailing their architectures, supported data sources, extensibility, and typical use cases for efficiently gathering and processing large‑scale data.

Big DataChukwaFluentd
0 likes · 11 min read
Common Big Data Collection Tools and Their Core Features
Smart Sea Tide
Smart Sea Tide
May 29, 2026 · Big Data

Designing a Scalable Big Data Service Platform Architecture

The article outlines how big data technology has evolved from core storage, processing, and analysis to include management, circulation, and security, forming a comprehensive ecosystem that now emphasizes cost reduction and enhanced security, and it details the platform's collection governance, analysis, visualization, and overall technical architecture.

Big Dataarchitecturedata analysis
0 likes · 3 min read
Designing a Scalable Big Data Service Platform Architecture
Alibaba Cloud Big Data AI Platform
Alibaba Cloud Big Data AI Platform
May 28, 2026 · Artificial Intelligence

From Assisted to Autonomous: How DataWorks Data Agent Revolutionizes Data Intelligence

DataWorks Data Agent advances from an assisted, code‑completion tool to a fully autonomous data‑intelligent agent, using a dual‑engine CLI/Claw architecture, unified runtime, open Skill ecosystem, and CPU‑GPU co‑optimization to automatically understand requirements, explore data, generate code, execute tasks, and deliver end‑to‑end results for developers and operators.

AIBig DataCLI
0 likes · 10 min read
From Assisted to Autonomous: How DataWorks Data Agent Revolutionizes Data Intelligence
DataFunSummit
DataFunSummit
May 28, 2026 · Artificial Intelligence

How DataWorks Data Agent Advances from Augmented Assistance to Full Autonomy

The article analyzes DataWorks Data Agent’s evolution from a helper‑style tool to an autonomous data‑centric AI agent, detailing its five‑stage roadmap, dual‑engine CLI/Claw architecture, unified runtime kernel, open skill ecosystem, and CPU‑GPU joint optimization for enterprise‑grade data automation.

AIBig DataData Agent
0 likes · 12 min read
How DataWorks Data Agent Advances from Augmented Assistance to Full Autonomy
DataFunTalk
DataFunTalk
May 28, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based ad‑hoc analysis to a Lambda‑style architecture and finally to a lakehouse with generic incremental compute, cutting architecture complexity, resource and development costs by one‑third while delivering second‑level queries over trillions of rows.

Big DataClickHouseFlink
0 likes · 21 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
Data Integration and Governance
Data Integration and Governance
May 27, 2026 · Big Data

10 Essential Data Cleaning Techniques Every AI Project Needs

The article outlines ten practical data‑cleaning methods—covering missing‑value imputation, duplicate handling, outlier detection, normalization, discretization, text cleaning, type conversion, multi‑source alignment, feature engineering, and sensitive‑data masking—explaining why each step matters for reliable AI model training.

Big Datadata cleaningdata masking
0 likes · 13 min read
10 Essential Data Cleaning Techniques Every AI Project Needs
DataFunTalk
DataFunTalk
May 25, 2026 · Big Data

MaxCompute’s AI‑Ready Evolution: Architecture, Features, and Real‑World Use Cases

This article examines how Alibaba Cloud’s MaxCompute platform has been transformed for AI workloads, detailing its multi‑layer architecture, multimodal data storage, SQL AI functions, the Python‑based MaxFrame framework, and real‑world deployments in large‑model preprocessing, autonomous driving, and multimodal image labeling.

AIBig DataDistributed Computing
0 likes · 12 min read
MaxCompute’s AI‑Ready Evolution: Architecture, Features, and Real‑World Use Cases
AI Large-Model Wave and Transformation Guide
AI Large-Model Wave and Transformation Guide
May 25, 2026 · Artificial Intelligence

AI‑Powered Underwater Simulation: Autonomous Perception, Decision & Execution

The article presents a comprehensive AI‑driven framework for unmanned underwater vehicles, detailing a three‑layer decision architecture, human‑machine collaboration models, conflict‑resolution mechanisms, data acquisition and simulation pipelines, ontology‑based knowledge graphs, and self‑evolution processes to enable reliable autonomous perception, planning, and actuation in complex marine environments.

Artificial IntelligenceBig DataR&D management
0 likes · 30 min read
AI‑Powered Underwater Simulation: Autonomous Perception, Decision & Execution
Big Data Tech Team
Big Data Tech Team
May 24, 2026 · Big Data

Data Warehouse Interview Pitfall Guide 2.0: Avoid Common SQL, Modeling, and ETL Mistakes

This guide compiles the most frequent interview pitfalls for data warehouse roles, covering SQL join and aggregation errors, window function misuse, subquery versus CTE performance myths, dimensional modeling mistakes, SCD implementation traps, layered design issues, data quality handling, ETL traps, Hive and Spark performance questions, real‑time warehousing considerations, and effective interview strategies.

Big DataETLHive
0 likes · 3 min read
Data Warehouse Interview Pitfall Guide 2.0: Avoid Common SQL, Modeling, and ETL Mistakes
DataFunTalk
DataFunTalk
May 22, 2026 · Big Data

How Xiaohongshu Cut Data Architecture Complexity and Cost by One‑Third in the Big AI Data Era

The article details Xiaohongshu's evolution from a simple ClickHouse‑based analytics layer to a Lambda‑enabled 2.0 stack and finally a Lakehouse‑based 3.0 architecture, showing how each iteration reduced infrastructure complexity, resource consumption and development effort by roughly one‑third while supporting trillions of daily events and AI‑driven use cases.

Big DataClickHouseFlink
0 likes · 21 min read
How Xiaohongshu Cut Data Architecture Complexity and Cost by One‑Third in the Big AI Data Era
DataFunSummit
DataFunSummit
May 21, 2026 · Big Data

Alibaba Cloud’s Agent-Ready Big Data AI Infrastructure: Boosting Data Development from Hours to Minutes

Facing a projected 85% of enterprises deploying internal agents within two years, Alibaba Cloud proposes an Agent-Ready big‑data AI infrastructure—comprising a unified data lake, real‑time processing, high‑dimensional vector retrieval, elastic model serving, and comprehensive security governance—that has already cut data‑development cycles from hours to 5‑10 minutes in internal model‑training and Taobao flash‑sale scenarios.

AIAgent-ReadyBig Data
0 likes · 15 min read
Alibaba Cloud’s Agent-Ready Big Data AI Infrastructure: Boosting Data Development from Hours to Minutes
DataFunSummit
DataFunSummit
May 20, 2026 · Big Data

How Kuaishou’s Real‑Time Data Lake Boosts AI and BI Architecture

The article explains how Kuaishou partnered with Apache Hudi to overhaul its ODS‑based data lake, addressing latency, storage cost, and complexity for AI and BI workloads, detailing the evolution from mysql‑to‑hive to mysql‑to‑hudi 1.0 and 2.0, the resulting performance gains, cost savings, and future roadmap.

AIBIBig Data
0 likes · 20 min read
How Kuaishou’s Real‑Time Data Lake Boosts AI and BI Architecture
Linyb Geek Road
Linyb Geek Road
May 20, 2026 · Big Data

Why 90% of Companies Get Data Governance Wrong and How to Reduce Friction

Most data‑governance initiatives fail not because of lacking technology but because they add friction; the article explains how companies mistakenly focus on rules, platforms, and processes, and offers a step‑by‑step approach—identifying high‑value tables, minimal metadata, targeted quality rules, and fast issue diagnosis—to make governance truly useful.

Big Datadata governancedata lineage
0 likes · 29 min read
Why 90% of Companies Get Data Governance Wrong and How to Reduce Friction
DataFunTalk
DataFunTalk
May 19, 2026 · Industry Insights

From Single‑Point Copilot to Platform‑Level Agentic: Real Challenges and Future Forks for Data Platforms

A live discussion dissected the shift from single‑point Copilot assistants to platform‑level Agentic data platforms, exposing hard architectural, security, knowledge‑base, evaluation, stability‑cost, and governance challenges while debating whether the future will favor a super‑agent or a multi‑agent ecosystem.

Big DataEnterprise GovernanceEvaluation
0 likes · 18 min read
From Single‑Point Copilot to Platform‑Level Agentic: Real Challenges and Future Forks for Data Platforms
Subtle Storm
Subtle Storm
May 18, 2026 · Fundamentals

Essential Architecture Exam Topics: A Must‑Read Review Guide

This guide compiles the most critical architecture concepts for the software architect certification, covering mandatory styles, quality‑attribute analysis, ATAM evaluation, microservice vs. SOA/monolith trade‑offs, Lambda/Kappa big‑data designs, cloud‑native fundamentals, high‑concurrency web patterns, distributed‑system theories, DDD, AI‑ops, IoT/edge computing, blockchain basics, and DevOps practices, each illustrated with concrete metrics and decision‑making steps.

AIOpsBig DataBlockchain
0 likes · 11 min read
Essential Architecture Exam Topics: A Must‑Read Review Guide
Data Integration and Governance
Data Integration and Governance
May 18, 2026 · Big Data

Why a Data Middle Platform Is Essential: A Complete Guide to Its Architecture and Components

The article explains why a data middle platform is a prerequisite for AI initiatives, outlines its functional architecture—data asset, tool platform, and application layers—and details the technical stack from ingestion and storage to compute, governance, and service layers, providing concrete examples and best‑practice recommendations.

Big DataETLdata architecture
0 likes · 14 min read
Why a Data Middle Platform Is Essential: A Complete Guide to Its Architecture and Components
DataFunSummit
DataFunSummit
May 17, 2026 · Industry Insights

From Single‑point Copilot to Platform‑level Agentic: Real Challenges and Future Paths for Data Platforms

A 90‑minute live discussion with data experts from vivo and YangQianGuan reveals that moving from a simple Copilot assistant to a platform‑level Agentic data system requires fundamental architectural changes, new infrastructure for memory, planning, tool orchestration, security guardrails, knowledge management, robust evaluation, and a clear ROI strategy.

AI governanceBig DataROI
0 likes · 19 min read
From Single‑point Copilot to Platform‑level Agentic: Real Challenges and Future Paths for Data Platforms
Data Party THU
Data Party THU
May 15, 2026 · Artificial Intelligence

2026 Big Data Challenge Announces Monthly Star Winners and Shares Winning Teams’ Insights

The 2026 China University Computer Competition – Big Data Challenge reveals the Monthly Star award winners, each receiving 800 RMB, and presents detailed experience reports from the top teams covering feature engineering, model selection, training validation, and ensemble strategies for stock prediction.

Big DataMachine LearningStock Prediction
0 likes · 7 min read
2026 Big Data Challenge Announces Monthly Star Winners and Shares Winning Teams’ Insights
dbaplus Community
dbaplus Community
May 14, 2026 · Big Data

Building a ‘One‑Sentence Bank’: Big Data and AI Fusion for Small Banks

The article outlines the evolution of big data in banking, compares management models for heterogeneous data, describes the shift from data engineering to knowledge engineering, introduces LLMOps for high‑quality knowledge bases, and details how integrating AI and data can enable a “one‑sentence bank” that answers queries and executes tasks.

Artificial IntelligenceBankingBig Data
0 likes · 22 min read
Building a ‘One‑Sentence Bank’: Big Data and AI Fusion for Small Banks
vivo Internet Technology
vivo Internet Technology
May 13, 2026 · Big Data

How Vivo Upgraded a Million‑Node YARN Cluster: Architecture, Scheduler Switch, and Performance Optimizations

This article details Vivo's end‑to‑end upgrade of a YARN 2.6.0 cluster to a modern version for a million‑node, hundred‑thousand‑tasks‑per‑day platform, covering architectural evolution, scheduler migration, compatibility fixes, performance tuning, and service‑continuity strategies.

Big DataCapacity SchedulerHadoop
0 likes · 28 min read
How Vivo Upgraded a Million‑Node YARN Cluster: Architecture, Scheduler Switch, and Performance Optimizations
DeWu Technology
DeWu Technology
May 13, 2026 · Big Data

How BP Claw Solves AI Coding Input Challenges in FlinkSpec’s Real‑Time Data Warehouse

The article explains how BP Claw tackles unstable AI coding results by automatically converting low‑quality PRD documents into structured, high‑quality requirements, applying token‑saving strategies, strict hallucination guards, and multi‑skill orchestration, which together boost FlinkSpec’s real‑time data‑warehouse delivery efficiency by up to 30%.

AI codingBP ClawBig Data
0 likes · 17 min read
How BP Claw Solves AI Coding Input Challenges in FlinkSpec’s Real‑Time Data Warehouse
DataFunTalk
DataFunTalk
May 11, 2026 · Big Data

How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based ad‑hoc analysis to a Lambda‑style architecture and finally to a lakehouse built on Iceberg, StarRocks, Flink and Spark, cutting architecture complexity, resource and development costs by two‑thirds while supporting trillions of daily events with sub‑second query latency.

Big DataClickHouseFlink
0 likes · 22 min read
How Xiaohongshu Re‑engineered Its Data Architecture for the Big AI Data Era
DataFunTalk
DataFunTalk
May 8, 2026 · Big Data

How MaxCompute Evolves into a Data+AI Platform: Architecture, Core Capabilities, and Real-World Cases

The article explains how Alibaba Cloud's MaxCompute has been transformed into a cloud‑native Data+AI platform, detailing its layered architecture, multimodal storage, model management, hybrid compute scheduling, SQL AI functions, the MaxFrame Python framework, and several enterprise case studies that demonstrate performance gains and flexible resource orchestration.

AI integrationBig DataData+AI
0 likes · 11 min read
How MaxCompute Evolves into a Data+AI Platform: Architecture, Core Capabilities, and Real-World Cases
DataFunTalk
DataFunTalk
May 6, 2026 · Big Data

How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era

The article details Xiaohongshu's four‑stage data‑platform evolution—from a simple ClickHouse ad‑hoc setup to a Lambda‑based 2.0 design and finally a lakehouse‑driven 3.0 architecture—highlighting the adoption of general incremental compute, cost‑reduction to one‑third, performance gains of up to ten‑fold, and the SPOT standards that guide the new system.

Big DataClickHouseFlink
0 likes · 21 min read
How Xiaohongshu Evolved Its Data Architecture for the Big AI Data Era
DataFunTalk
DataFunTalk
Apr 29, 2026 · Big Data

How Xiaohongshu Revamped Its Data Architecture for the Big AI Data Era

Xiaohongshu transformed its data platform from a simple ClickHouse‑based analytics stack to a unified lakehouse with generic incremental compute, cutting architecture complexity, resource cost, and development effort by roughly one‑third while supporting petabyte‑scale, sub‑second queries across its 350 million‑user app.

Big DataClickHouseFlink
0 likes · 22 min read
How Xiaohongshu Revamped Its Data Architecture for the Big AI Data Era
Model Perspective
Model Perspective
Apr 28, 2026 · Big Data

How a Taiwan Ban Became Free Advertising for Amap’s Map App

A recent Taiwan government warning against Amap turned into a viral boost, exposing the app’s superior traffic‑light countdown, massive data‑driven network effects, and the underlying reverse‑propagation model that explains why the ban accelerated downloads rather than suppressing them.

AmapBig DataNetwork Effects
0 likes · 11 min read
How a Taiwan Ban Became Free Advertising for Amap’s Map App
DataFunTalk
DataFunTalk
Apr 28, 2026 · Artificial Intelligence

From “Lobster” to Ontology: DACon Reveals the Next Trend in Self‑Evolving AI Agents

The DACon conference in Shanghai gathered over 8,000 developers and experts, showcasing 50 talks that explored self‑evolving AI agents, the open‑source GenericAgent framework, data‑governance ontology, Agent‑Ready big‑data infrastructure, and AI+AR ecosystems, while highlighting practical case studies and future industry directions.

AI AgentsAI+ARBig Data
0 likes · 11 min read
From “Lobster” to Ontology: DACon Reveals the Next Trend in Self‑Evolving AI Agents
Subtle Storm
Subtle Storm
Apr 28, 2026 · Big Data

Big Data’s Commercial Value in E‑Commerce & Retail: From Insight to Decision

The article examines how big data moves retail beyond merely spotting trends to actively shaping pricing, inventory, real‑time monitoring, and strategic decisions, illustrating dynamic pricing, intelligent stock management, and live dashboards with examples from airlines, hotels, and Double‑11 sales.

Big DataReal‑time Monitoringdynamic pricing
0 likes · 5 min read
Big Data’s Commercial Value in E‑Commerce & Retail: From Insight to Decision
Subtle Storm
Subtle Storm
Apr 27, 2026 · Big Data

When Algorithms Learn to Read People: Big Data’s Commercial Value in E‑Commerce (Part 1)

The article explains how big‑data techniques such as collaborative filtering, time‑aware recommendations, and behavior‑sequence profiling transform e‑commerce and retail by revealing individual consumer intent, boosting conversion rates by up to 30% and enabling more precise pricing and inventory decisions.

Big DataRecommendation Systemsbehavior analysis
0 likes · 5 min read
When Algorithms Learn to Read People: Big Data’s Commercial Value in E‑Commerce (Part 1)
DataFunSummit
DataFunSummit
Apr 27, 2026 · Artificial Intelligence

How Tencent Games Leverages AI to Turn Data Governance into a Service

Tencent Games’ data governance team details an AI‑driven, end‑to‑end semantic framework that shifts traditional rule‑based data management to a service‑oriented model, cutting storage waste by 30 %, halving development time, and boosting asset recommendation accuracy to 95 % across its global gaming platform.

AIBig DataGaming Industry
0 likes · 19 min read
How Tencent Games Leverages AI to Turn Data Governance into a Service
DataFunSummit
DataFunSummit
Apr 25, 2026 · Big Data

AI‑Era Multimodal Data Lake Infrastructure: TBDS Design, Storage, Compute, and Governance

The article analyzes how Tencent Cloud's TBDS platform tackles the AI era's multimodal data lake challenges through a native storage format (Lance), elastic Ray‑based compute, standardized metadata with Gravitino, and automated governance via Lakekeeper, citing architecture details, performance numbers, and real‑world deployments.

AI infrastructureBig DataGravitino
0 likes · 13 min read
AI‑Era Multimodal Data Lake Infrastructure: TBDS Design, Storage, Compute, and Governance
DataFunSummit
DataFunSummit
Apr 24, 2026 · Artificial Intelligence

AI‑Driven Data Governance as a Service: Tencent Games' Paradigm Shift

This talk details how Tencent Games leverages AI to transform its data governance from rule‑based, passive processes into a semantic, service‑oriented paradigm, addressing resource waste, low collaboration efficiency, and scalability challenges while delivering measurable improvements in cost, speed, and asset quality.

AIBig DataGaming
0 likes · 19 min read
AI‑Driven Data Governance as a Service: Tencent Games' Paradigm Shift
DataFunTalk
DataFunTalk
Apr 22, 2026 · Industry Insights

How Xiaohongshu Cut Data Platform Costs by Two‑Thirds with Incremental Computing

This article details Xiaohongshu's journey from a ClickHouse‑based batch analytics stack to a unified lakehouse architecture powered by generic incremental computing, showing how the company reduced architecture complexity, resource consumption and development effort each to roughly one‑third while supporting trillions of daily events with sub‑10‑second query latency.

Big DataLakehouseXiaohongshu
0 likes · 24 min read
How Xiaohongshu Cut Data Platform Costs by Two‑Thirds with Incremental Computing
Big Data Tech Team
Big Data Tech Team
Apr 22, 2026 · Big Data

Inside Big Tech: Full Breakdown of AI Agents for Data Warehouse Governance

The article analyzes how leading internet companies embed AI agents across the entire data‑warehouse lifecycle to automate governance, presenting real‑world case studies from Alibaba, ByteDance, JD.com and Tencent, and quantifies benefits such as over 65% reduction in manual effort, 50% drop in metric duplication, and a 40% boost in resource utilization.

AI AgentsBig Dataautomation
0 likes · 10 min read
Inside Big Tech: Full Breakdown of AI Agents for Data Warehouse Governance
DataFunSummit
DataFunSummit
Apr 21, 2026 · Industry Insights

How SelectDB Cuts 60% Costs and Boosts Real‑Time Performance for New Energy Batteries

The whitepaper analyzes the data‑driven transformation of the new‑energy battery sector, outlines four core challenges—massive data streams, fast‑changing R&D demands, long manufacturing cycles, and multi‑dimensional quality standards—and demonstrates how SelectDB’s unified lake‑warehouse architecture delivers million‑level throughput, second‑level latency, up to 30× query speedup, and 60% cost reduction across real‑world case studies.

Big DataNew EnergyReal-Time Analytics
0 likes · 18 min read
How SelectDB Cuts 60% Costs and Boosts Real‑Time Performance for New Energy Batteries
DataFunSummit
DataFunSummit
Apr 19, 2026 · Big Data

How OPPO Built a Multi‑Modal Data Lake with Gravitino and Curvine

OPPO’s data‑lake team, led by David, detailed their transition from Hive‑Spark to a unified multi‑modal lake, leveraging Gravitino for cross‑engine metadata management and the open‑source Curvine cache to eliminate data silos, boost I/O performance, and support massive image, recommendation, and AI‑Agent workloads.

Big DataData LakeMultimodal
0 likes · 11 min read
How OPPO Built a Multi‑Modal Data Lake with Gravitino and Curvine
Big Data Tech Team
Big Data Tech Team
Apr 17, 2026 · Industry Insights

Can AI Replace Data Warehouse Engineers? Exploring the Future of Data Modeling

The article examines how large‑language‑model AI can automate data‑warehouse modeling tasks—generating SQL, designing schemas, handling ETL, and tracing lineage—while highlighting current pain points, practical limitations, and four emerging trends that will reshape the role of data engineers over the next few years.

AIBig DataLarge Language Models
0 likes · 11 min read
Can AI Replace Data Warehouse Engineers? Exploring the Future of Data Modeling
Ctrip Technology
Ctrip Technology
Apr 16, 2026 · Big Data

How Ray + DuckDB Cut 9B-Row Attribution Queries from 40s to 15s

When attribution analysis on over 900 million rows slowed to more than 40 seconds and threatened cluster stability, Ctrip's smart attribution team rebuilt the architecture with Ray and DuckDB, achieving sub‑15‑second query times, 160 % performance gain, and complete resource isolation.

Attribution AnalysisBig DataDistributed Computing
0 likes · 22 min read
How Ray + DuckDB Cut 9B-Row Attribution Queries from 40s to 15s
DataFunTalk
DataFunTalk
Apr 16, 2026 · Big Data

How Xiaohongshu Cut Data Architecture Costs by Two‑Thirds with Incremental Computing

This article details Xiaohongshu's data platform evolution from a simple ClickHouse‑based ad‑hoc system to a Lambda‑style architecture and finally a lakehouse solution, highlighting how the adoption of a new incremental computing model reduced architectural complexity, resource consumption and development effort each to roughly one‑third while delivering sub‑second query performance on petabyte‑scale data.

Big DataLakehouseXiaohongshu
0 likes · 21 min read
How Xiaohongshu Cut Data Architecture Costs by Two‑Thirds with Incremental Computing
DataFunSummit
DataFunSummit
Apr 15, 2026 · Industry Insights

Why Traditional Data Platforms Fail and How Ontology Drives Triple‑Digit ROI

The article analyzes costly data‑platform failures—such as a $40 million payroll system in San Francisco schools and a collapsed Healthcare.gov launch—identifies the root cause as ineffective data middle platforms, and demonstrates how Palantir’s ontology‑based three‑layer architecture (semantic, dynamics, decision) can turn data into actionable insights, delivering triple‑digit ROI for enterprises like BP, Novartis, and General Mills.

Big DataOntologyPalantir
0 likes · 5 min read
Why Traditional Data Platforms Fail and How Ontology Drives Triple‑Digit ROI
Data Integration and Governance
Data Integration and Governance
Apr 15, 2026 · Big Data

What Business Data Mining Actually Uncovers: Patterns, Relationships, Anomalies, and Trends

The article explains that enterprise data mining isn’t about extracting raw numbers but about discovering underlying business patterns, relationships, anomalies, and trends—such as churn risk, production issues, product bundling opportunities, and profit‑draining steps—while emphasizing that stable, integrated data foundations are the real prerequisite for valuable insights.

Big DataData Integrationbusiness analytics
0 likes · 10 min read
What Business Data Mining Actually Uncovers: Patterns, Relationships, Anomalies, and Trends
DataFunTalk
DataFunTalk
Apr 11, 2026 · Industry Insights

Why Most Intelligent Data Analytics Fail and How Aloudata’s Agent Architecture Solves It

This article examines three common misconceptions in enterprise intelligent data analysis, explains how a semantic metric layer can break data silos, and details Aloudata Agent’s dual‑path engine, multi‑agent collaboration, and product design that together deliver trustworthy, deep, and democratized analytics for modern businesses.

AIAttribution AnalysisBig Data
0 likes · 18 min read
Why Most Intelligent Data Analytics Fail and How Aloudata’s Agent Architecture Solves It
DataFunTalk
DataFunTalk
Apr 10, 2026 · Big Data

How Xiaohongshu Cut Data Architecture Costs by Two‑Thirds with Incremental Computing

This article analyzes Xiaohongshu's data platform evolution—from a simple ClickHouse‑based analytics layer to a Lambda architecture and finally a lakehouse design—highlighting how adopting a new incremental computing model reduced architecture complexity, resource consumption, and development effort each to roughly one‑third while delivering sub‑second query performance on petabyte‑scale data.

Big DataLakehouseXiaohongshu
0 likes · 22 min read
How Xiaohongshu Cut Data Architecture Costs by Two‑Thirds with Incremental Computing
Big Data Tech Team
Big Data Tech Team
Apr 9, 2026 · Industry Insights

Why Data Engineers Are the New AI Powerhouses: 4 Core Reasons & Actionable Tips

The article analyzes why data development engineers are becoming more valuable in the AI era, outlining four core reasons—including data‑driven AI limits, the rise of RAG architectures, heightened data compliance, and a talent shortage—while offering concrete advice on mastering real‑time pipelines, unstructured data, and AI infrastructure.

AI infrastructureBig DataRAG
0 likes · 8 min read
Why Data Engineers Are the New AI Powerhouses: 4 Core Reasons & Actionable Tips
Data Integration and Governance
Data Integration and Governance
Apr 8, 2026 · Big Data

Master the Four Data Integration Patterns in One Guide

This article explains the four common data integration patterns—ETL, ELT, API‑based, and message‑queue approaches—detailing their core workflows, suitable scenarios, advantages, and trade‑offs so readers can choose the method that best fits their business and technical constraints.

APIBig DataELT
0 likes · 10 min read
Master the Four Data Integration Patterns in One Guide
Alibaba Cloud Observability
Alibaba Cloud Observability
Apr 6, 2026 · Cloud Native

How Alibaba Cloud Built Real‑Time OpenAPI Monitoring with Flink + SLS

This article details the design and implementation of a cloud‑native, real‑time monitoring system for Alibaba Cloud OpenAPI, covering background challenges, a Flink‑SLS architecture, multi‑region data processing, checkpoint and state‑backend tuning, source‑side predicate pushdown, visualization with Grafana, and production results.

Big DataFlinkPredicate Pushdown
0 likes · 21 min read
How Alibaba Cloud Built Real‑Time OpenAPI Monitoring with Flink + SLS
Big Data Tech Team
Big Data Tech Team
Apr 1, 2026 · Big Data

Why Your 2026 Big Data Resume Is Being Ignored and How to Fix It

In the 2026 spring hiring season, many big‑data job seekers see their resumes disappear because they still focus on offline batch processing, while employers now demand real‑time streaming, AI‑driven data pipelines, and cloud‑native deployment skills such as Flink, vector databases, and Kubernetes.

AI integrationBig DataFlink
0 likes · 7 min read
Why Your 2026 Big Data Resume Is Being Ignored and How to Fix It