Tagged articles

data architecture

248 articles · Page 3 of 3
Programmer DD
Programmer DD
Dec 11, 2019 · Big Data

Big Data Architecture Secrets: Storage-Compute Separation & Spark in Action

This article explores how enterprises can tackle the explosive growth of data by adopting modern big‑data architectures, including storage‑compute separation, data‑driven workflows, risk‑control frameworks, and real‑world Spark optimizations, offering practical guidance for scalable, high‑performance analytics.

Big DataSparkStorage Compute Separation
0 likes · 12 min read
Big Data Architecture Secrets: Storage-Compute Separation & Spark in Action
UCloud Tech
UCloud Tech
Dec 4, 2019 · Big Data

How to Evolve Big Data Architectures for ZB‑Scale Analytics and Real‑World Use Cases

This article reviews the challenges of handling Zettabyte‑scale data, outlines practical big‑data processing architectures, discusses storage‑compute separation, data‑driven workflows, risk‑control frameworks, and shares concrete Spark implementations at MobTech, offering actionable insights for modern data engineers.

SparkStorage Compute Separationdata architecture
0 likes · 13 min read
How to Evolve Big Data Architectures for ZB‑Scale Analytics and Real‑World Use Cases
21CTO
21CTO
Nov 27, 2019 · Big Data

How Xiaohongshu Scales Real‑Time Personalized Recommendations with Flink

The article summarizes Guo Yi’s 2019 Alibaba Cloud conference talk, outlining Xiaohongshu’s personalized recommendation architecture, detailing the data stack from ingestion to warehouse, and showcasing a Flink‑based real‑time multi‑dimensional user behavior aggregation use case, followed by a vision for the next year’s data architecture evolution.

FlinkReal-time Streamingdata architecture
0 likes · 3 min read
How Xiaohongshu Scales Real‑Time Personalized Recommendations with Flink
Architecture Digest
Architecture Digest
Nov 5, 2019 · Big Data

Architecture Overview of Taobao, Meituan, and Didi Big Data Platforms

This article examines the big‑data architectures of three leading Chinese internet companies—Taobao, Meituan, and Didi—detailing their data sources, synchronization mechanisms, batch and streaming processing layers, and the common scheduling components that unify their Hadoop‑based ecosystems.

Big DataDidiHadoop
0 likes · 7 min read
Architecture Overview of Taobao, Meituan, and Didi Big Data Platforms
Big Data Technology & Architecture
Big Data Technology & Architecture
Oct 28, 2019 · Big Data

Big Data Technology and Architecture: Leveraging Spark and HBase for Real‑Time and Offline Processing

This article outlines the challenges of various big‑data scenarios such as financial risk control, recommendation systems, and social feeds, explains why Spark is chosen over alternatives, describes a one‑stop data platform architecture with Spark‑HBase integration, and shares best‑practice tips and case studies.

Big DataHBaseSpark
0 likes · 7 min read
Big Data Technology and Architecture: Leveraging Spark and HBase for Real‑Time and Offline Processing
dbaplus Community
dbaplus Community
Oct 27, 2019 · Product Management

What Skills Do Data Product Managers Need in a Data Middle Platform?

The article explains the concept of a data middle platform, why it matters for rapid demand response and resource integration, and outlines the distinct responsibilities and required skill sets of data product managers and data platform product managers within such ecosystems.

Data GovernanceData Productdata architecture
0 likes · 11 min read
What Skills Do Data Product Managers Need in a Data Middle Platform?
dbaplus Community
dbaplus Community
Oct 22, 2019 · Big Data

How Weibo Built a Billion‑Log Real‑Time Data Platform with Flink

This article details how Weibo’s advertising team designed and implemented a real‑time data platform capable of processing over a hundred billion daily logs, covering technology selection, Flink advantages, architecture evolution, data processing pipelines, component libraries, fault‑tolerance strategies, and the construction of a multi‑layer real‑time data warehouse.

Big DataCheckpointData Warehouse
0 likes · 25 min read
How Weibo Built a Billion‑Log Real‑Time Data Platform with Flink
Architects' Tech Alliance
Architects' Tech Alliance
Oct 17, 2019 · Big Data

Understanding Alibaba's Data Middle Platform: Concepts, Architecture, and Differences from Data Warehouses and Data Lakes

The article explains Alibaba's data middle platform—its definition, methodology, organizational structure, key tools, and how it differs from traditional data warehouses and data lakes—while highlighting its role in supporting scalable, business‑centric data services and digital transformation.

AlibabaBig DataData Lake
0 likes · 16 min read
Understanding Alibaba's Data Middle Platform: Concepts, Architecture, and Differences from Data Warehouses and Data Lakes
Meituan Technology Team
Meituan Technology Team
Oct 17, 2019 · Big Data

OneData Methodology: Building a Unified Data Warehouse Architecture and Governance Framework

By adapting Alibaba’s OneData methodology, the project establishes a unified data‑warehouse architecture, standards, and governance framework—including consolidated business intake, standardized design layers, naming conventions, and delivery metrics—that resolves data‑quality issues, enhances scalability and reusability, and delivers faster, reliable data support for evolving business needs.

Big DataData GovernanceData Modeling
0 likes · 15 min read
OneData Methodology: Building a Unified Data Warehouse Architecture and Governance Framework
iQIYI Technical Product Team
iQIYI Technical Product Team
Sep 12, 2019 · Big Data

iQIYI's Big Data Architecture Evolution and Adoption of Druid

iQIYI upgraded its big‑data stack by adopting Druid as the core engine for free‑time queries and ElasticSearch for pre‑computed fixed‑time queries, overcoming early API, security and scaling challenges through monthly segment granularity, parallel sub‑queries, Redis caching and failover, cutting typical query latency from over two seconds to about 150 ms and reaching 99.9 % service success.

Bitmap IndexElasticsearchdata architecture
0 likes · 12 min read
iQIYI's Big Data Architecture Evolution and Adoption of Druid
Alibaba Cloud Developer
Alibaba Cloud Developer
Sep 4, 2019 · Big Data

How Structured Big Data Storage Powers Modern Data Systems

This article explores the core components of data systems, the evolution toward lightweight, intelligent big data architectures, the distinction between primary and secondary storage, challenges of data replication, and how Alibaba Cloud's Tablestore implements advanced features such as storage‑compute separation, CDC, and multi‑model indexing for scalable, cost‑effective structured big data storage.

Big DataCDCcloud services
0 likes · 24 min read
How Structured Big Data Storage Powers Modern Data Systems
37 Interactive Technology Team
37 Interactive Technology Team
Mar 28, 2019 · Big Data

Approaches to Building a Basic Data Platform

To handle terabytes of daily data and diverse business needs, the company built a three‑layer basic data platform—collection/computation/storage, unified data management, and API‑driven services—augmented by a standardized collection system, a robust Domino scheduler, and a self‑service analysis tool, aiming to evolve into a full data‑middle‑office for end‑to‑end intelligence.

Data IntegrationSchedulingSelf‑service analytics
0 likes · 8 min read
Approaches to Building a Basic Data Platform
DataFunTalk
DataFunTalk
Feb 18, 2019 · Big Data

Hulu’s Big Data Architecture and Sophon OLAP Cache Layer Overview

This article presents an in‑depth overview of Hulu’s big‑data platform, detailing its multi‑layer architecture, the design and functionality of the Sophon OLAP cache layer, and how Impala is employed for high‑performance query processing and integration with cloud‑native engines.

HuluImpalaSophon
0 likes · 16 min read
Hulu’s Big Data Architecture and Sophon OLAP Cache Layer Overview
JD Tech
JD Tech
Jan 28, 2019 · Big Data

Technical Overview of JD Marketing 360: 4A Consumer Asset Model, 4E Marketing Framework, and Big Data Architecture

The article presents a comprehensive technical analysis of JD's Marketing 360 project, detailing the three industry pain points, the 4A consumer‑asset model, the 4E marketing methodology, and the underlying big‑data, AI‑driven architecture that enables real‑time analytics, multi‑modal model training, and performance optimizations.

Artificial IntelligenceCustomer Asset ModelMTA
0 likes · 14 min read
Technical Overview of JD Marketing 360: 4A Consumer Asset Model, 4E Marketing Framework, and Big Data Architecture
ITPUB
ITPUB
Oct 23, 2018 · Big Data

How Meituan Built a Scalable Real‑Time Data Warehouse with Flink

This article explains how Meituan tackled growing real‑time data demands by redesigning its streaming platform, adopting a layered real‑time data warehouse architecture, selecting storage and compute technologies such as Cellar, Elasticsearch, Druid and Flink, and sharing practical tips on dimension expansion, joins, and aggregation to achieve higher throughput and lower latency.

FlinkMeituanReal-time Data Warehouse
0 likes · 15 min read
How Meituan Built a Scalable Real‑Time Data Warehouse with Flink
DataFunTalk
DataFunTalk
Sep 9, 2018 · Big Data

Druid Principles and Their Application in Insurance Data Analytics

This article summarizes a presentation by Ping An Insurance data engineers on Druid’s architecture, core concepts, node roles, tuning strategies, and real-world deployment for insurance analytics, illustrating how Druid enables sub‑second, high‑cardinality OLAP queries and supports both real‑time and batch processing.

DruidInsurancedata architecture
0 likes · 11 min read
Druid Principles and Their Application in Insurance Data Analytics
Big Data and Microservices
Big Data and Microservices
Sep 4, 2018 · Big Data

Exploring Five Big Data Architectures—from Traditional to Unified AI Designs

The article examines the evolution of big‑data processing by comparing five prevalent architectures—traditional Hadoop‑based stacks, streaming‑only designs, Kappa, Lambda, and the unified Unifield model—highlighting their strengths, weaknesses, and suitable scenarios while discussing the limitations of classic BI systems and the role of distributed storage, computation, and machine‑learning integration.

Big DataHadoopKappa
0 likes · 14 min read
Exploring Five Big Data Architectures—from Traditional to Unified AI Designs
Architects' Tech Alliance
Architects' Tech Alliance
Jul 27, 2018 · Backend Development

Designing Data Architecture for Microservices: Principles, Patterns, and Database Choices

This article explains how to design data architecture for microservice systems, covering microservice fundamentals, advantages, decoupling, lightweight APIs, continuous delivery, database per service versus shared databases, polyglot persistence, scaling dimensions, sharding strategies, and why MongoDB is a suitable choice.

MongoDBPolyglot PersistenceSharding
0 likes · 15 min read
Designing Data Architecture for Microservices: Principles, Patterns, and Database Choices
58 Tech
58 Tech
Jun 27, 2018 · Big Data

Overview of the 58 User Profile System Architecture and Data Processing

The article describes the design, data integration, ID mapping, tag generation, and application scenarios of the 58 user profiling platform, which aggregates billions of user IDs across multiple business lines to provide online and offline persona data for personalization, analytics, and AI modeling.

Big DataData IntegrationID mapping
0 likes · 12 min read
Overview of the 58 User Profile System Architecture and Data Processing
High Availability Architecture
High Availability Architecture
May 21, 2018 · Big Data

Interview with Baidu’s Chief Big Data Architect Ma Ruyue on OLAP, HTAP, and Emerging Big Data Technologies

In this interview, Baidu’s senior big‑data architect Ma Ruyue discusses his career transition from Hadoop to online databases, the design philosophy behind Baidu’s Palo ROLAP system, the future of HTAP, and his views on the evolving big‑data ecosystem including Spark, AI, and containerization.

Distributed ComputingHTAPdata architecture
0 likes · 11 min read
Interview with Baidu’s Chief Big Data Architect Ma Ruyue on OLAP, HTAP, and Emerging Big Data Technologies
Architecture Digest
Architecture Digest
Mar 27, 2018 · Backend Development

Data Architecture Design in Microservice Development

This article explains the multi‑layer data architecture design for microservice systems, covering concepts such as data usability, primary and secondary data decoupling, sharding, multi‑source data adaptation and caching, and introduces data marts to improve scalability and maintainability.

Data CachingData MartMicroservices
0 likes · 10 min read
Data Architecture Design in Microservice Development
Alibaba Cloud Developer
Alibaba Cloud Developer
Nov 3, 2017 · Big Data

How Alibaba Built an EB-Scale, Real-Time Big Data Platform

Alibaba’s senior data expert Yao Bin Hui explains how the company constructed a standardized, end-to-end big-data ecosystem—from low-level data collection and AI algorithms to data services and product platforms—enabling petabyte-scale integration and second-level response times that power both internal operations and millions of external users.

AlibabaBig Datadata architecture
0 likes · 10 min read
How Alibaba Built an EB-Scale, Real-Time Big Data Platform
21CTO
21CTO
Jul 3, 2017 · Big Data

Inside the World’s Best Data Architectures: Netflix, Facebook, Airbnb, Pinterest

This article explores the cutting‑edge data pipelines of Netflix, Facebook, Airbnb and Pinterest, detailing the massive event volumes they handle, the core technologies such as Kafka, Spark, Presto and Hadoop, and how these giants design scalable, real‑time analytics infrastructures.

AirbnbBig DataFacebook
0 likes · 6 min read
Inside the World’s Best Data Architectures: Netflix, Facebook, Airbnb, Pinterest
dbaplus Community
dbaplus Community
Jan 8, 2017 · Big Data

How to Build a Cost‑Effective Data Platform for Small‑to‑Medium Enterprises

This article explains why data platforms are essential for modern SMEs, defines what a data platform is, outlines a four‑step methodology (source definition, analysis theme, ETL processing, and reporting), and shares architectural choices, team structures, common pitfalls, and practical advice for rapid, iterative implementation.

Data WarehouseETLSME
0 likes · 15 min read
How to Build a Cost‑Effective Data Platform for Small‑to‑Medium Enterprises
Tencent Cloud Developer
Tencent Cloud Developer
Jan 6, 2017 · Game Development

Challenges and Design Considerations for Game Server Data Systems

Game server development suffers from generic client‑communication tools and inadequate data stores, leading to duplicated, latency‑heavy code, so a purpose‑built, memory‑resident distributed cache that persists locally and eliminates serialization boiler‑plate is essential for real‑time, low‑latency gameplay.

cachingdata architecturegame server
0 likes · 12 min read
Challenges and Design Considerations for Game Server Data Systems
Liulishuo Tech Team
Liulishuo Tech Team
Dec 16, 2016 · Cloud Computing

Key Takeaways from AWS re:Invent 2016: New Services, Data Architecture, and Operational Insights

The article shares a comprehensive recap of AWS re:Invent 2016, highlighting over a thousand new features across compute, storage, networking, security and tools, discussing serverless concepts, the Modern Data Architecture framework, and practical lessons learned by the English‑Fluent‑Talk engineering team.

AWSCloud Computingdata architecture
0 likes · 11 min read
Key Takeaways from AWS re:Invent 2016: New Services, Data Architecture, and Operational Insights
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Dec 13, 2016 · Big Data

Umeng’s Mobile Big Data Platform: Architecture, Challenges & Insights

The article details Umeng’s mobile big‑data platform architecture, describing its Lambda‑style hybrid design, data ingestion pipeline with dual Kafka clusters, offline and real‑time processing using Hadoop, Spark, Storm, and storage layers such as HDFS, HBase, MongoDB and Elasticsearch, while also discussing challenges in data collection, cleaning, computation, security, and value‑added services.

HadoopKafkaLambda Architecture
0 likes · 13 min read
Umeng’s Mobile Big Data Platform: Architecture, Challenges & Insights
21CTO
21CTO
Apr 14, 2016 · Big Data

How Meituan’s Data Architecture Powers Precise Mobile Marketing

This article details Meituan Dianping's data‑driven approach to precise marketing, describing the O2O marketing framework, a layered pyramid data system, profiling techniques, budget monitoring, and two real‑world case studies that together illustrate how big‑data technologies boost marketing efficiency on mobile platforms.

Big Datadata architecturemachine learning
0 likes · 12 min read
How Meituan’s Data Architecture Powers Precise Mobile Marketing
Meituan Technology Team
Meituan Technology Team
Apr 14, 2016 · Big Data

Data‑Driven Precise Marketing: Architecture and Case Studies at Meituan‑Dianping

Meituan‑Dianping’s data‑driven precise‑marketing platform combines a layered pyramid architecture—data warehouse, service, and front‑end layers—with real‑time profile services powered by Redis and Elasticsearch, offering tools such as Hoek, Cord, and Cloud/Star to automate audience selection, coupon recommendation, and KPI monitoring, illustrated by food‑delivery user discovery and WeChat red‑packet coupon case studies, and guided by principles of reusable models and SOA decoupling.

Case studydata architectureprecise marketing
0 likes · 9 min read
Data‑Driven Precise Marketing: Architecture and Case Studies at Meituan‑Dianping
Architecture Digest
Architecture Digest
Apr 14, 2016 · Big Data

Data‑Driven Precise Marketing: Architecture and Case Studies from Meituan Dianping

This article presents Meituan Dianping's data‑driven precise marketing architecture, detailing a layered pyramid system, user profiling, budget monitoring, and two real‑world cases—potential user mining and a smart coupon engine—demonstrating how big‑data techniques improve marketing efficiency and ROI.

Meituandata architecturemachine learning
0 likes · 12 min read
Data‑Driven Precise Marketing: Architecture and Case Studies from Meituan Dianping
21CTO
21CTO
Mar 31, 2016 · Big Data

Inside Airbnb’s Massive Big Data Platform: Architecture, Lessons & Scaling Secrets

Airbnb’s engineering team outlines the evolution of its big‑data platform, detailing the philosophy behind its architecture, the dual “gold” and “silver” Hive clusters, migration to Mesos, use of Presto, Airpal, Airflow, and the performance and cost gains achieved through these design choices.

AirbnbAirflowBig Data
0 likes · 11 min read
Inside Airbnb’s Massive Big Data Platform: Architecture, Lessons & Scaling Secrets
21CTO
21CTO
Feb 12, 2016 · Backend Development

Key Challenges in Building High‑Traffic Data‑Intensive Web Platforms

This article examines the critical issues of massive data handling, concurrency, file storage, relational design, indexing, distributed processing, AJAX usage, security, clustering, and OpenAPI trends that developers must address when architecting large, high‑interaction web sites.

backenddata architecturedistributed systems
0 likes · 8 min read
Key Challenges in Building High‑Traffic Data‑Intensive Web Platforms
21CTO
21CTO
Jan 28, 2016 · Databases

How LinkedIn Scales Data: Inside Its Multi‑Phase Database Architecture

This article explains how LinkedIn manages real‑time profile updates, news feeds, and social graph data through a three‑phase architecture that combines RDBMS, NoSQL stores, caching, and Lucene indexes to achieve high consistency, availability, and partition tolerance at massive scale.

LinkedInNoSQLdata architecture
0 likes · 11 min read
How LinkedIn Scales Data: Inside Its Multi‑Phase Database Architecture
ITPUB
ITPUB
Jan 20, 2016 · Big Data

How Meizu Built an Agile Big Data Platform for Millions of Users

The Meizu Tech Open Day showcased the company's rapid evolution to a data‑driven mobile internet firm, detailing its DW1.0 and DW2.0 data‑warehouse architectures, recommendation pipelines, Spark adoption, and ELK‑based log analytics, while sharing practical lessons and future challenges.

Big DataData WarehouseELK
0 likes · 11 min read
How Meizu Built an Agile Big Data Platform for Millions of Users
21CTO
21CTO
Dec 30, 2015 · Big Data

Mastering Massive Data: MapReduce, Hadoop, and Taobao’s Architecture

This article introduces the fundamental MapReduce model and Hadoop framework, explains their roles in large‑scale data processing, and then examines Taobao’s massive‑data product architecture—including its data source, compute, storage, query, and product layers, as well as the MyFOX, Prom, and Glider components and caching strategies.

HadoopMapReduceTaobao
0 likes · 16 min read
Mastering Massive Data: MapReduce, Hadoop, and Taobao’s Architecture
21CTO
21CTO
Dec 9, 2015 · Big Data

Mastering Hadoop: From MapReduce Basics to Taobao’s Massive Data Architecture

This article introduces the fundamental MapReduce model and Hadoop framework, explains their components such as HDFS, MapReduce, and HBase, and then examines Taobao’s large‑scale data product architecture—including storage, computation, query, and caching layers—to illustrate practical big‑data processing techniques.

HadoopMapReducedata architecture
0 likes · 17 min read
Mastering Hadoop: From MapReduce Basics to Taobao’s Massive Data Architecture
Architect
Architect
Dec 2, 2015 · Big Data

Designing an Agile Data Warehouse Architecture for Internet Companies

The article outlines a practical, end‑to‑end data platform architecture for internet businesses, covering data collection, storage and analysis, sharing, real‑time processing, task scheduling, and the importance of simplicity and agility in building an agile data warehouse.

Big DataData WarehouseHadoop
0 likes · 10 min read
Designing an Agile Data Warehouse Architecture for Internet Companies
21CTO
21CTO
Nov 26, 2015 · Big Data

Understanding Big Data: 4V Traits, Google’s Distributed Computing, and Hadoop Ecosystem

This article explores the 4V characteristics of big data, real‑world data growth examples, historical analogies, Google’s GFS‑MapReduce‑BigTable model, Hadoop’s architecture and HDFS processes, HBase components, NoSQL alternatives, and practical big‑data applications at Tencent and beyond.

Distributed ComputingHadoopMapReduce
0 likes · 7 min read
Understanding Big Data: 4V Traits, Google’s Distributed Computing, and Hadoop Ecosystem
Suning Technology
Suning Technology
May 22, 2015 · Big Data

Suning’s Big Data Platform Evolution: From SAP BW to Real‑Time Streaming

This article chronicles Suning’s journey from early SAP‑based data warehouses to a modern, open‑source big data platform featuring real‑time collection, Hadoop‑Hive offline processing, Storm‑based streaming, and a visual development environment, highlighting how each layer addresses growing data volume, variety, and business demands.

HadoopStormdata architecture
0 likes · 5 min read
Suning’s Big Data Platform Evolution: From SAP BW to Real‑Time Streaming
MaGe Linux Operations
MaGe Linux Operations
Feb 25, 2015 · Big Data

Do You Really Need Hadoop? 10 Alternatives to Consider First

This article explains why many companies over‑invest in Hadoop, outlines how to evaluate data size, growth, and relevance, and presents practical alternatives such as archiving, data sampling, database sharding, and hiring business‑savvy analysts before committing to a Hadoop deployment.

Big Data AlternativesData ArchivingHadoop
0 likes · 9 min read
Do You Really Need Hadoop? 10 Alternatives to Consider First