Tagged articles

distributed systems

2274 articles · Page 20 of 23
Tencent Cloud Developer
Tencent Cloud Developer
Jun 1, 2018 · Backend Development

Building Tencent Xinge: Architecture and Practices for Massive Mobile Push Service

The talk details Tencent Xinge’s architecture and cloud‑native practices that enable hundred‑billion‑level mobile push, combining terminal integration, real‑time backend filtering, distributed bitmap selection, precise‑push AI models, and DevOps pipelines to deliver fast, scalable, data‑driven notifications with effect tracking.

Real-time Analyticsbackend architecturebig data
0 likes · 18 min read
Building Tencent Xinge: Architecture and Practices for Massive Mobile Push Service
Tencent Cloud Developer
Tencent Cloud Developer
May 31, 2018 · Backend Development

Tencent Billing System (Mi Master): Architecture, Reliability, Security, and Global Capabilities

The Mi Master billing platform, a SaaS service from Tencent that processes over 735 billion RMB quarterly, provides a unified, modular architecture with distributed multi‑master databases, high‑availability cross‑region deployment, multi‑stage fraud detection, and global support for 80+ payment channels across 180+ countries, delivering seamless APIs, automated reconciliation, and extensive revenue‑sharing tools for products such as Honor of Kings, PUBG, WeChat Pay, and QQ Wallet.

billing architecturedistributed systemspayment system
0 likes · 19 min read
Tencent Billing System (Mi Master): Architecture, Reliability, Security, and Global Capabilities
Efficient Ops
Efficient Ops
May 30, 2018 · Databases

How SF Express Transformed Its Database Operations: From Legacy to Open‑Source, Distributed, and Intelligent Ops

This talk details SF Express’s journey from heterogeneous legacy databases to standardized open‑source, distributed architectures and intelligent operations, covering standardization, migration to open‑source, scaling with Mycat, automated resource pooling, and the ThinkDB platform that drives proactive, automated DBA workflows.

Mycatautomationdatabase
0 likes · 18 min read
How SF Express Transformed Its Database Operations: From Legacy to Open‑Source, Distributed, and Intelligent Ops
Efficient Ops
Efficient Ops
May 21, 2018 · Databases

Why Do Database Failures Happen and How to Prevent Them?

This article examines common hardware and network failures in data centers, analyzes real‑world outage cases, classifies fault domains, and presents comprehensive strategies for database fault handling—including logging, checkpointing, backup, replication, and high‑availability architectures—to improve reliability and reduce downtime.

backupdatabasedistributed systems
0 likes · 22 min read
Why Do Database Failures Happen and How to Prevent Them?
Java Backend Technology
Java Backend Technology
May 20, 2018 · Backend Development

Which Cache Update Strategy Guarantees Consistency? A Deep Dive into DB‑Cache Synchronization

This article examines three common cache‑update approaches—updating the cache after the database, deleting the cache before updating the database, and updating the database then deleting the cache—analyzes their drawbacks, and presents practical solutions such as delayed double‑delete and retry mechanisms to ensure data consistency.

BackendCache invalidationConsistency
0 likes · 10 min read
Which Cache Update Strategy Guarantees Consistency? A Deep Dive into DB‑Cache Synchronization
Alibaba Cloud Developer
Alibaba Cloud Developer
May 16, 2018 · Cloud Computing

From Mall to Cloud: How Alibaba’s Tech Evolution Shaped Modern Cloud Computing

Senior Alibaba engineer Xiao Xie recounts his decade‑long journey from the early Taobao Mall project to leading the Cloud Computing “Flying‑Sky Eight” team, detailing pivotal initiatives like the Five‑Color Stone integration, full‑link stress testing for Double‑11, and the evolution toward self‑developed, distributed cloud technologies.

AlibabaCloud ComputingTech Interview
0 likes · 12 min read
From Mall to Cloud: How Alibaba’s Tech Evolution Shaped Modern Cloud Computing
ITFLY8 Architecture Home
ITFLY8 Architecture Home
May 12, 2018 · Backend Development

What Drives the Architecture of Billion‑User Platforms? Lessons from Weibo

This article explores the essence of system architecture for massive web services, illustrating strategic and tactical considerations through examples like Uber and Weibo, and discusses key capabilities such as abstraction, classification, performance, service decomposition, multi‑level caching, distributed tracing, and continuous learning for scalable backend design.

backend designdistributed systemslarge-scale web
0 likes · 21 min read
What Drives the Architecture of Billion‑User Platforms? Lessons from Weibo
Qunar Tech Salon
Qunar Tech Salon
May 11, 2018 · Databases

Minsheng Bank’s Distributed Transformation and NewSQL Practice with SequoiaDB

The article details Minsheng Bank’s shift to distributed architecture, outlining regulatory drivers, business requirements, the adoption of sharding, cross‑center high‑availability, and new‑type distributed databases, and showcases performance results of SequoiaDB 3.0 across multiple high‑throughput banking scenarios.

NewSQLSequoiaDBbanking
0 likes · 9 min read
Minsheng Bank’s Distributed Transformation and NewSQL Practice with SequoiaDB
21CTO
21CTO
May 9, 2018 · Operations

How Alipay Built Seamless High Availability and Disaster Recovery for Millions of Transactions

This article examines Alipay's evolution from a simple single‑datacenter setup to a multi‑active‑active, unit‑based architecture, detailing the technical challenges of high availability, disaster recovery, failover design, blue‑green deployment, and how these solutions enable continuous service during massive traffic spikes like Double 11.

AlipayBlue-Green Deploymentdisaster recovery
0 likes · 17 min read
How Alipay Built Seamless High Availability and Disaster Recovery for Millions of Transactions
Architecture Digest
Architecture Digest
May 9, 2018 · Operations

High Availability and Disaster Recovery Architecture: The Evolution of Alipay’s System Design

This article examines the importance of high‑availability and disaster‑recovery architectures, tracing Alipay’s evolution from a simple load‑balanced setup through multi‑datacenter, failover, and unit‑based designs that address scalability, data consistency, and continuous service delivery challenges.

disaster recoverydistributed systemsfailover
0 likes · 16 min read
High Availability and Disaster Recovery Architecture: The Evolution of Alipay’s System Design
21CTO
21CTO
May 5, 2018 · Backend Development

From Single Server to Scalable Architecture: Key Lessons from Large‑Scale Site Design

This comprehensive note distills the evolution of large‑website architecture—from single‑server setups to layered, distributed, and highly available systems—covering caching, clustering, read/write separation, CDN, NoSQL, business splitting, scalability, extensibility, and automation strategies.

distributed systemshigh availabilitylarge-scale architecture
0 likes · 20 min read
From Single Server to Scalable Architecture: Key Lessons from Large‑Scale Site Design
Architecture Digest
Architecture Digest
May 5, 2018 · Backend Development

Evolution and Core Principles of Large‑Scale Website Architecture

This article summarizes the evolution stages, architectural patterns, and key concerns such as performance, scalability, extensibility, high availability, and distributed design that large‑scale websites must address, providing practical insights and visual diagrams for each concept.

Cachingdistributed systemshigh availability
0 likes · 21 min read
Evolution and Core Principles of Large‑Scale Website Architecture
Mike Chen's Internet Architecture
Mike Chen's Internet Architecture
May 3, 2018 · Backend Development

Fundamentals and Evolution of Large-Scale Website Architecture Design

This article explains the essence of software architecture as a process of reducing system entropy through splitting and merging, outlines the capabilities required of architects, and details the step‑by‑step evolution of large‑scale website infrastructures including caching, CDN, database sharding, and messaging systems.

database shardingdistributed systemsscalability
0 likes · 10 min read
Fundamentals and Evolution of Large-Scale Website Architecture Design
ITFLY8 Architecture Home
ITFLY8 Architecture Home
May 2, 2018 · Backend Development

How WeChat’s SeqSvr Generates Trillions of Sequence Numbers Daily

This article explains the design and evolution of WeChat's high‑availability seqsvr service, which provides per‑user 64‑bit sequence numbers for data synchronization, handling trillions of requests with millisecond latency through pre‑allocation, section sharing, and layered storage architecture.

distributed systemsscalabilitysequence generation
0 likes · 10 min read
How WeChat’s SeqSvr Generates Trillions of Sequence Numbers Daily
Architecture Digest
Architecture Digest
Apr 29, 2018 · Backend Development

Designing High‑Concurrency Architecture for Large‑Scale E‑Commerce Applications

This article outlines practical strategies for building high‑concurrency back‑end systems—including server architecture, load balancing, database clustering, caching, message queues, asynchronous processing, and service‑oriented design—to ensure smooth operation of traffic‑intensive e‑commerce services.

CachingHigh Concurrencybackend architecture
0 likes · 19 min read
Designing High‑Concurrency Architecture for Large‑Scale E‑Commerce Applications
Java Captain
Java Captain
Apr 26, 2018 · Backend Development

Dubbo Overview, Architecture, and a Step‑by‑Step Demo with Zookeeper and Spring

This article introduces Dubbo’s background, explains the evolution of e‑commerce architectures to RPC‑based distributed systems, details Dubbo’s components, advantages, and drawbacks, and provides a complete Maven‑based demo—including Zookeeper installation, Spring configuration, and Java code—for building and consuming a Dubbo service.

DubboJavaRPC
0 likes · 19 min read
Dubbo Overview, Architecture, and a Step‑by‑Step Demo with Zookeeper and Spring
Meituan Technology Team
Meituan Technology Team
Apr 19, 2018 · Backend Development

How Meituan Waimai Supports Ten Million Daily Orders: Evolution of Its Backend Architecture

Meituan Waimai handles ten‑million daily orders by evolving from a tiny monolithic prototype to a distributed, micro‑service‑based platform that uses sharded databases, caches, set‑based traffic partitioning, automated AIOps, dynamic container scaling, prioritized degradation switches, and AI‑driven features to sustain massive, growing traffic.

High ConcurrencyMeituanWaimai
0 likes · 19 min read
How Meituan Waimai Supports Ten Million Daily Orders: Evolution of Its Backend Architecture
Architecture Digest
Architecture Digest
Apr 18, 2018 · Databases

Understanding Distributed Architecture and Its Applications in MySQL and Large‑Scale Systems

The article explains the concept of distributed architecture, its key characteristics such as cohesion and transparency, showcases how MySQL and middleware like Mycat are used in e‑commerce platforms, and outlines the evolution, practical implementations, and challenges of building scalable distributed database systems.

MySQLMycatbig data
0 likes · 15 min read
Understanding Distributed Architecture and Its Applications in MySQL and Large‑Scale Systems
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Apr 11, 2018 · Fundamentals

Mastering Distributed System Design: Core Principles Every Engineer Should Know

This article outlines essential distributed system concepts—including system decomposition, concurrency, caching strategies, online vs. offline processing, push/pull communication, load limiting, service degradation, CAP theorem, and eventual consistency—to help engineers design scalable, reliable architectures for high‑traffic applications.

CAP theoremdistributed systemsmicroservices
0 likes · 13 min read
Mastering Distributed System Design: Core Principles Every Engineer Should Know
Architecture Digest
Architecture Digest
Apr 10, 2018 · Fundamentals

Reliability, Scalability, and Maintainability in Distributed System Design

This article examines core distributed system design principles—reliability, scalability, and maintainability—explaining how techniques such as replication, partitioning, consensus algorithms, and transactions address hardware, software, and human failures, and discusses vertical and horizontal scaling strategies to achieve robust, extensible, and maintainable architectures.

ConsensusReplicationdistributed systems
0 likes · 8 min read
Reliability, Scalability, and Maintainability in Distributed System Design
dbaplus Community
dbaplus Community
Apr 8, 2018 · Databases

Mastering Multi‑Tenant Load Balancing in Alibaba Cloud Table Store

This article explains the architecture, data model, and multi‑tenant load‑balancing strategies of Alibaba Cloud Table Store, detailing the challenges of distributed NoSQL systems and presenting practical solutions for resource quantification, fairness, trigger timing, and SLA‑driven automation.

Alibaba CloudNoSQLTable Store
0 likes · 20 min read
Mastering Multi‑Tenant Load Balancing in Alibaba Cloud Table Store
21CTO
21CTO
Apr 6, 2018 · Cloud Native

Why Service Mesh Is the Next Evolution of Microservices

This article examines the limitations of traditional microservice frameworks, introduces service mesh as a solution with sidecar architecture, outlines its definition, evolution stages, and timeline, and concludes with resources for further learning and practical implementation.

cloud nativedistributed systemsmicroservices
0 likes · 9 min read
Why Service Mesh Is the Next Evolution of Microservices
Architecture Digest
Architecture Digest
Apr 6, 2018 · Cloud Native

An Overview of Service Mesh: Addressing the Limitations of Traditional Microservices

This article reviews the challenges of early microservice frameworks—high technical barriers, limited multi‑language support, and intrusive code—and explains how service mesh architectures with sidecar proxies, exemplified by Linkerd, Envoy, and Istio, provide a dedicated, language‑agnostic infrastructure layer that simplifies service governance and operations.

Istiocloud nativedistributed systems
0 likes · 9 min read
An Overview of Service Mesh: Addressing the Limitations of Traditional Microservices
Efficient Ops
Efficient Ops
Apr 1, 2018 · Backend Development

Ele.me’s Secret to Seamless Multi-Region Active-Active Architecture

This article details how Ele.me engineered a cross‑region active‑active system that scales elastically, tolerates whole‑data‑center failures, and maintains real‑time food‑delivery performance through geographic sharding, intelligent routing, and robust data‑replication middleware.

data replicationdistributed systemsgeographic sharding
0 likes · 18 min read
Ele.me’s Secret to Seamless Multi-Region Active-Active Architecture
AntTech
AntTech
Mar 29, 2018 · Artificial Intelligence

Ant Group CTO Cheng Li’s Money 20/20 Asia Presentation on FinTech Innovation: AI, Blockchain, Cloud and Mobile Payments

In his Money 20/20 Asia keynote, Ant Group CTO Cheng Li outlines the company’s fintech roadmap, highlighting AI‑driven risk engines, blockchain‑based trust mechanisms, cloud‑native infrastructure, and innovative mobile payment solutions that aim to make financial services more inclusive and efficient.

Artificial IntelligenceCloud ComputingFinTech
0 likes · 16 min read
Ant Group CTO Cheng Li’s Money 20/20 Asia Presentation on FinTech Innovation: AI, Blockchain, Cloud and Mobile Payments
Architecture Digest
Architecture Digest
Mar 28, 2018 · Operations

Implementing High-Concurrency Performance Testing and Practical Solutions Based on Server Architecture

This article explains the concept of high concurrency, outlines a server architecture that supports it—including load balancing, distributed databases, NoSQL caches and CDN—and presents practical testing methods and implementation patterns such as caching strategies and message‑queue designs to handle massive simultaneous requests.

CachingHigh ConcurrencyServer Architecture
0 likes · 7 min read
Implementing High-Concurrency Performance Testing and Practical Solutions Based on Server Architecture
Architecture Digest
Architecture Digest
Mar 26, 2018 · Operations

Alipay’s Double 11 Architecture: Logical Data Centers, Distributed Transactions, and High‑Availability Strategies

The article details Alipay’s comprehensive architecture for the Double 11 shopping festival, covering its three‑layer IAAS/PAAS/SAAS model, logical data‑center design, multi‑active disaster‑recovery, blue‑green deployment, distributed data sharding, transaction processing, and the Ant Credit Pay service’s performance and risk‑control mechanisms.

Alipayarchitecturebig data
0 likes · 16 min read
Alipay’s Double 11 Architecture: Logical Data Centers, Distributed Transactions, and High‑Availability Strategies
Architecture Digest
Architecture Digest
Mar 20, 2018 · Backend Development

Source Code Analysis, Distributed Architecture, Microservices, Performance Optimization, and Java Engineering Overview

This article discusses the importance of source code analysis, outlines key concepts in distributed systems, explains microservice architecture, highlights performance optimization techniques for Java applications, and presents practical engineering advice for modern backend development.

Java engineeringSource Code Analysisdistributed systems
0 likes · 8 min read
Source Code Analysis, Distributed Architecture, Microservices, Performance Optimization, and Java Engineering Overview
Java Backend Technology
Java Backend Technology
Mar 19, 2018 · Fundamentals

Why Distributed Consistency Matters: From CAP to BASE Explained

This article explores the importance of data consistency in distributed systems, illustrating real‑world scenarios, explaining consistency models such as strong, weak and eventual, and detailing the challenges and theories like CAP and BASE that guide system designers in balancing consistency, availability, and partition tolerance.

BASE theoryCAP theoremConsistency
0 likes · 18 min read
Why Distributed Consistency Matters: From CAP to BASE Explained
Efficient Ops
Efficient Ops
Mar 15, 2018 · Operations

How Baidu’s CCS System Scales Command Execution Across Millions of Servers

This article examines Baidu’s Cluster Control System (CCS), detailing its two‑level data model, four‑tier scheduling architecture, and three‑layer execution agents, and explains how control and execution information, redundancy, and fault‑tolerant designs enable reliable large‑scale command execution across thousands of servers.

Command Executiondistributed systemsoperations
0 likes · 12 min read
How Baidu’s CCS System Scales Command Execution Across Millions of Servers
Efficient Ops
Efficient Ops
Mar 15, 2018 · Operations

Mastering Large-Scale Command Execution: From Basics to Baidu’s Cluster Control System

This article explores the fundamentals of command execution, examines the challenges of scaling command delivery across hundreds of thousands of servers, and details Baidu’s Cluster Control System architecture that enables efficient, flexible, and extensible distributed command management for operations teams.

Command Executiondeploymentdistributed systems
0 likes · 10 min read
Mastering Large-Scale Command Execution: From Basics to Baidu’s Cluster Control System
Java Backend Technology
Java Backend Technology
Mar 13, 2018 · Fundamentals

Why Consistent Hashing Is the Key to Scalable Redis Clusters

This article explains the limitations of simple modulo hashing for Redis clusters, introduces consistent hashing with a virtual‑node ring to achieve fault tolerance and seamless scaling, and demonstrates how the algorithm reduces data skew and improves cache performance in distributed systems.

Cachingconsistent hashingdistributed systems
0 likes · 11 min read
Why Consistent Hashing Is the Key to Scalable Redis Clusters
360 Zhihui Cloud Developer
360 Zhihui Cloud Developer
Mar 1, 2018 · Operations

AI-Driven Strategies for Optimizing Resource Management in Distributed Systems

This article reviews cloud gaming resource management, introduces search‑engine instance distribution techniques, explores AI‑based disk‑failure prediction and load forecasting, and presents replica and DDoS‑detection strategies to improve efficiency and reliability of large‑scale distributed systems.

AIdistributed systemsfailure prediction
0 likes · 12 min read
AI-Driven Strategies for Optimizing Resource Management in Distributed Systems
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Feb 28, 2018 · Backend Development

Inside Alibaba’s Live Streaming Architecture: Lessons from a Senior Engineer

In this extensive interview, senior Alibaba engineer Chen Kangxian shares his experiences designing large‑scale distributed systems, live‑streaming platforms, and high‑concurrency architectures, offering practical insights on technology choices, failure handling, and career growth for software architects.

High Concurrencydistributed systemslive streaming
0 likes · 34 min read
Inside Alibaba’s Live Streaming Architecture: Lessons from a Senior Engineer
Hulu Beijing
Hulu Beijing
Feb 28, 2018 · Big Data

How Hulu’s Nesto Engine Delivers Near‑Real‑Time OLAP on TB‑Scale Data

This article introduces Hulu's in‑house OLAP engine Nesto, detailing its near‑real‑time data ingestion, nested data model, TB‑level storage using HBase and Parquet, MPP query execution, custom predicate library, and the overall architecture that enables sub‑second ad‑hoc queries for user analytics.

HBaseOLAPQuery Engine
0 likes · 22 min read
How Hulu’s Nesto Engine Delivers Near‑Real‑Time OLAP on TB‑Scale Data
Java Backend Technology
Java Backend Technology
Feb 27, 2018 · Backend Development

Mastering Large-Scale Website Architecture: 10 Essential Patterns Explained

This article outlines ten fundamental architecture patterns for high‑traffic websites—including layering, partitioning, distribution, clustering, caching, asynchronous processing, redundancy, automation, and security—explaining their goals, benefits, challenges, and best‑practice constraints to help engineers build scalable, reliable, and maintainable systems.

Backend PatternsCachingautomation
0 likes · 11 min read
Mastering Large-Scale Website Architecture: 10 Essential Patterns Explained
Java Backend Technology
Java Backend Technology
Feb 22, 2018 · Backend Development

From Single Server to Global Scale: Evolution of Large Website Architecture

This article explores the defining traits of large‑scale websites and walks through the step‑by‑step evolution of their architecture—from single‑server setups to distributed systems with caching, load balancing, database sharding, and micro‑services—while highlighting common design pitfalls and best‑practice recommendations.

Cachingbackend architecturedistributed systems
0 likes · 8 min read
From Single Server to Global Scale: Evolution of Large Website Architecture
AI Cyberspace
AI Cyberspace
Jan 29, 2018 · Backend Development

Mastering Celery: Periodic Tasks, Sync Calls, Result Storage, and Monitoring

Explore how to configure Celery’s periodic (Beat) tasks, perform synchronous task calls, persist results using Redis, monitor workers with Flower, and debug remotely via telnet, with practical code examples and step‑by‑step instructions for robust backend task management.

PythonTask Queuebackend development
0 likes · 7 min read
Mastering Celery: Periodic Tasks, Sync Calls, Result Storage, and Monitoring
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jan 28, 2018 · Backend Development

Designing Scalable E‑Commerce Architecture: From Business to Technical Layers

This article explores how to design a high‑performance, highly available, and scalable e‑commerce platform by separating business and technical architectures, detailing subsystem decomposition, scaling strategies, and the evolution from simple single‑server setups to distributed, clustered solutions.

Backendarchitecturedistributed systems
0 likes · 13 min read
Designing Scalable E‑Commerce Architecture: From Business to Technical Layers
Meituan Technology Team
Meituan Technology Team
Jan 26, 2018 · Big Data

Design and Implementation of a Real-Time Data Processing System at Meituan

Meituan designed a Storm‑based real‑time data processing platform that guarantees at‑least‑once delivery and high availability, employs a custom spout, regression‑driven traffic smoothing, and a low‑latency KV store with atomic operations, persisting results in Kafka, MySQL and Cellar to power merchant dashboards and heat‑tag analytics, while planning broader real‑time analytics expansion.

Real-time DataStormbig data
0 likes · 10 min read
Design and Implementation of a Real-Time Data Processing System at Meituan
Java Backend Technology
Java Backend Technology
Jan 21, 2018 · Backend Development

Is Microservices Doomed? Uncovering the Hidden Complexities Behind the Hype

The article critically examines micro‑services, outlining their promised benefits such as independent development, deployment and scaling, while exposing the hidden operational, dev‑ops, state‑management, communication, versioning and distributed‑transaction challenges that can turn them into a fragile, overly complex system.

Complexitydistributed systemsmicroservices
0 likes · 15 min read
Is Microservices Doomed? Uncovering the Hidden Complexities Behind the Hype
Architect's Tech Stack
Architect's Tech Stack
Jan 20, 2018 · Backend Development

What Is Microservices? Concepts, Design Guidelines, Integration Patterns, and Trade‑offs

This article explains the microservices architecture style, compares it with monolithic and SOA approaches, outlines design principles, communication mechanisms, data decentralization, integration patterns, and discusses the advantages and disadvantages of adopting microservices in modern software systems.

Service Integrationdesign principlesdistributed systems
0 likes · 16 min read
What Is Microservices? Concepts, Design Guidelines, Integration Patterns, and Trade‑offs
Vipshop Quality Engineering
Vipshop Quality Engineering
Jan 17, 2018 · Backend Development

Why Zookeeper Connections Fail After 1 MB and How to Fix Them

A staging environment’s new scheduled task kept failing due to Zookeeper disconnections caused by packets exceeding the default 1 MB maxBuffer, and the article explains the root cause, heartbeat timing, and how adjusting Djute.maxbuffer or upgrading Zookeeper resolves the issue.

BackendTroubleshootingZooKeeper
0 likes · 4 min read
Why Zookeeper Connections Fail After 1 MB and How to Fix Them
21CTO
21CTO
Jan 16, 2018 · Fundamentals

Why Distributed Consensus Is So Hard: From CAP to Byzantine Fault Tolerance

Distributed systems rely on consensus to ensure consistent results, but achieving it faces fundamental challenges such as network unreliability, node failures, and trade‑offs captured by the CAP theorem, FLP impossibility, and various algorithms like Paxos, Raft, and Byzantine Fault Tolerance, each balancing consistency, availability, and safety.

Byzantine Fault ToleranceCAP theoremPaxos
0 likes · 26 min read
Why Distributed Consensus Is So Hard: From CAP to Byzantine Fault Tolerance
Architecture Digest
Architecture Digest
Jan 16, 2018 · Fundamentals

Consistency, Consensus, and Reliability in Distributed Systems

This article explains the core challenges of achieving consistency in distributed systems, describes consensus algorithms such as Paxos and Raft, discusses theoretical limits like the FLP impossibility and CAP theorem, and shows how trade‑offs among consistency, availability, and partition tolerance shape practical system design.

CAP theoremConsistencyFLP impossibility
0 likes · 24 min read
Consistency, Consensus, and Reliability in Distributed Systems
iQIYI Technical Product Team
iQIYI Technical Product Team
Jan 12, 2018 · Backend Development

Microservice Implementation Experience in iQIYI Bubble Backend System

Facing over 60 million daily users and 100 K QPS, iQIYI’s Bubble platform migrated from a monolithic codebase to a business‑driven microservice architecture—splitting services by entity and function, adopting the internal RPCHUB RPC framework, establishing ownership, fault‑tolerance, monitoring and CI/CD pipelines, and addressing scaling challenges to sustain rapid growth.

RPCarchitecturedistributed systems
0 likes · 20 min read
Microservice Implementation Experience in iQIYI Bubble Backend System
DevOps
DevOps
Jan 9, 2018 · Fundamentals

Git Basics for Enterprise Developers – Core Concepts and Advantages

This article introduces the fundamental concepts of Git for enterprise developers, covering its distributed version control model, core operations such as commits, branches, and file states, and highlighting its advantages like parallel development, faster releases, strong community support, and integration with modern tooling.

CommitGitbranching
0 likes · 12 min read
Git Basics for Enterprise Developers – Core Concepts and Advantages
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Jan 7, 2018 · Operations

Mastering Load Balancing: Algorithms, Code Samples, and Real‑World Insights

This article explains the concept of load balancing in distributed systems, outlines its benefits for throughput and reliability, compares common architectural layers, evaluates key algorithmic considerations, and provides Python implementations of round‑robin, weighted, random, hash‑based, and least‑connection strategies along with deployment options.

NetworkingPythonalgorithm
0 likes · 14 min read
Mastering Load Balancing: Algorithms, Code Samples, and Real‑World Insights
Java Backend Technology
Java Backend Technology
Jan 2, 2018 · Operations

When to Adopt Distributed Architecture? 5 Common Patterns Explained

This article explains why and when to move to distributed architecture, outlines the typical upgrade and splitting steps, and details five common distributed cluster patterns—including load balancing, leader election, blockchain, master‑slave, and consistent hashing—highlighting their trade‑offs and use cases.

Leader Electionarchitecture patternsconsistent hashing
0 likes · 8 min read
When to Adopt Distributed Architecture? 5 Common Patterns Explained
Alibaba Cloud Developer
Alibaba Cloud Developer
Dec 28, 2017 · Backend Development

Explore Alibaba’s 2017 Open‑Source Powerhouse: Dubbo, RocketMQ, Ant Design & More

The article reviews Alibaba’s 2017 open‑source contributions, highlighting nine major projects—including Dubbo, RocketMQ, Druid, Fastjson, ApsaraCache, Pouch, Dragonfly, Ant Design, and Egg—detailing their features, community impact, and how they advance high‑performance distributed systems, cloud‑native computing, and modern application development.

Alibababackend developmentdistributed systems
0 likes · 23 min read
Explore Alibaba’s 2017 Open‑Source Powerhouse: Dubbo, RocketMQ, Ant Design & More
Architecture Digest
Architecture Digest
Dec 27, 2017 · Backend Development

Handling Transactions, Failover, and Exactly‑Once Semantics in Distributed Systems

This article explores how distributed systems determine node liveness, manage failover and recovery, and implement at‑most‑once, at‑least‑once, and exactly‑once processing guarantees—including opaque transactions and two‑phase commit—using examples from Kafka, Zookeeper, and big‑data pipelines.

Exactly-OnceZooKeeperbig data
0 likes · 15 min read
Handling Transactions, Failover, and Exactly‑Once Semantics in Distributed Systems
Alibaba Cloud Developer
Alibaba Cloud Developer
Dec 27, 2017 · Databases

How Alibaba’s Next‑Gen Database Powered Double 11: Elasticity, Cloud & AI

Alibaba’s database team explains how their next‑generation X‑DB system achieved extreme elasticity, high performance, and cost efficiency during the Double 11 shopping festival by leveraging cloud‑native hybrid deployment, containerization, storage‑compute separation, Paxos‑based consistency, and AI‑driven self‑optimizing DBA tools, while outlining key challenges and solutions.

AI OpsCloud Computingdatabase
0 likes · 12 min read
How Alibaba’s Next‑Gen Database Powered Double 11: Elasticity, Cloud & AI
Java Backend Technology
Java Backend Technology
Dec 18, 2017 · Backend Development

How to Scale Websites for Massive Data and High Concurrency

This article outlines practical strategies for building and scaling web applications—covering caching, static page generation, database optimization, read/write separation, NoSQL, Hadoop, distributed deployment, service separation, and CDN—to handle massive data volumes and high‑traffic loads efficiently.

BackendCachingHigh Concurrency
0 likes · 21 min read
How to Scale Websites for Massive Data and High Concurrency
Architecture Digest
Architecture Digest
Dec 15, 2017 · Backend Development

Evolution and Practice of Suning E‑commerce Inventory System Architecture for Double 11 Peak

This article details the business scope, challenges, architectural evolution, and practical solutions of Suning's inventory system—including front‑mid‑back separation, self‑developed high‑concurrency services, unitization, multi‑active deployment, and pre‑Double 11 capacity planning—to ensure stable, scalable e‑commerce operations during massive traffic spikes.

Double 11High ConcurrencyInventory
0 likes · 17 min read
Evolution and Practice of Suning E‑commerce Inventory System Architecture for Double 11 Peak
Efficient Ops
Efficient Ops
Dec 11, 2017 · Databases

Inside Twitter’s Manhattan: How a Massive Distributed Database Powers Real‑Time Ads

The article explores Twitter’s Manhattan storage system, detailing its architecture, CAP trade‑offs across various database types, the design of its modular storage engines, high‑performance operations, and the DevOps practices that enable reliable, low‑latency handling of billions of requests in a massive distributed environment.

CAP theoremDevOpsTwitter
0 likes · 14 min read
Inside Twitter’s Manhattan: How a Massive Distributed Database Powers Real‑Time Ads
Architecture Digest
Architecture Digest
Dec 9, 2017 · Databases

Understanding Relational and Distributed Transactions

This article explains the fundamentals of relational database transactions, the ACID properties, and various distributed transaction protocols such as 2PC, 3PC, TCC, message‑based approaches, and 1PC, while discussing their advantages, drawbacks, and practical considerations.

2PCACIDTCC
0 likes · 20 min read
Understanding Relational and Distributed Transactions
21CTO
21CTO
Dec 6, 2017 · Backend Development

How Ctrip Built a High‑Performance Distributed ID Generator for Millions of Users

This article explains Ctrip's design of a globally unique, high‑concurrency user ID generator for sharded MySQL databases, reviews common industry solutions, and details the final optimized approach that combines a MySQL auto‑increment table with in‑memory segment allocation to achieve millisecond‑level response times.

High ConcurrencyJavaMySQL
0 likes · 8 min read
How Ctrip Built a High‑Performance Distributed ID Generator for Millions of Users
Efficient Ops
Efficient Ops
Nov 27, 2017 · Operations

How Facebook Scales to Billions: Disaggregated Networks, Storage, and Warm Spark

Facebook’s journey from early startup ops to supporting over 2 billion monthly users reveals how disaggregated network, storage, and warm‑storage‑enabled Spark architectures overcome scalability bottlenecks, illustrating the operational strategies and design principles that power massive, reliable data‑center services.

Cloud Infrastructurebig datadistributed systems
0 likes · 12 min read
How Facebook Scales to Billions: Disaggregated Networks, Storage, and Warm Spark
dbaplus Community
dbaplus Community
Nov 26, 2017 · Databases

Understanding HBase Region Auto‑Splitting: Policies, Process, and Pitfalls

This article explains how HBase achieves scalable region auto‑splitting, detailing the various split policies, the algorithm for locating split points, the transactional split workflow, reference file handling, data migration via compaction, cleanup procedures, and common troubleshooting tips.

HBaseReference FileRegion Split
0 likes · 17 min read
Understanding HBase Region Auto‑Splitting: Policies, Process, and Pitfalls
21CTO
21CTO
Nov 26, 2017 · Backend Development

How to Implement Effective Rate Limiting with Guava and Redis

This article explains why rate limiting is essential for high‑traffic services, describes the token‑bucket algorithm, shows how to use Guava's RateLimiter and Cache for single‑node limits, and presents a Redis‑based solution that works across distributed instances.

BackendGuavaRedis
0 likes · 7 min read
How to Implement Effective Rate Limiting with Guava and Redis
MaGe Linux Operations
MaGe Linux Operations
Nov 25, 2017 · Fundamentals

Why Do Open‑Source Gurus Still Rely on IRC? The Surprising Reasons

This article explains what IRC is and outlines eight technical advantages—distributed architecture, low bandwidth, privacy, cross‑platform support, and a rich bot ecosystem—plus additional cultural reasons why many open‑source developers continue to favor this decades‑old chat protocol.

IRCdistributed systemslegacy tools
0 likes · 4 min read
Why Do Open‑Source Gurus Still Rely on IRC? The Surprising Reasons
Yuewen Technology
Yuewen Technology
Nov 24, 2017 · Backend Development

How to Build a Scalable Distributed Task Scheduler from Scratch

This article outlines the shortcomings of using crontab for large‑scale job execution, defines the requirements for a custom distributed scheduler, describes its three‑component architecture (trigger, monitor, management), and details key technical solutions such as process isolation, distributed locking, and log aggregation.

QuartzRediscron replacement
0 likes · 12 min read
How to Build a Scalable Distributed Task Scheduler from Scratch
21CTO
21CTO
Nov 21, 2017 · Operations

How We Scaled WeChat Pay’s Transaction Records to Billions Daily

This article chronicles the evolution of WeChat Pay’s transaction record system—from early key/value storage bottlenecks and incomplete data to a distributed, tiered architecture that supports billions of daily records, improves query performance, ensures data security, and handles holiday traffic spikes through flexible throttling.

WeChat Paydata securitydistributed systems
0 likes · 11 min read
How We Scaled WeChat Pay’s Transaction Records to Billions Daily
21CTO
21CTO
Nov 21, 2017 · Backend Development

How Uber Scales Its Real-Time Ride‑Sharing Platform: Architecture Secrets

This article examines Uber's rapid 38‑fold growth by detailing the design, scaling techniques, and fault‑tolerance mechanisms of its real‑time market platform, including geographic indexing, microservices, distributed storage, and the DISCO scheduling system.

Uberdistributed systemsreal-time platform
0 likes · 19 min read
How Uber Scales Its Real-Time Ride‑Sharing Platform: Architecture Secrets
Architecture Digest
Architecture Digest
Nov 20, 2017 · Cloud Native

Evolution of Microservice Systems Toward a Reactive Microsystem Architecture

The article explains how traditional microservice architectures evolve into event‑driven reactive microsystems by adopting events‑first DDD, reactive design, and event‑based persistence, highlighting the role of the Actor model, asynchronous non‑blocking communication, event sourcing, and saga‑based distributed transaction handling.

DDDEvent Sourcingdistributed systems
0 likes · 9 min read
Evolution of Microservice Systems Toward a Reactive Microsystem Architecture
MaGe Linux Operations
MaGe Linux Operations
Nov 20, 2017 · Backend Development

Mastering Web Crawlers: Core Principles, Architecture, and Modern Challenges

This article explains how web crawlers work—from initial URL seeding and request handling to flow control, content extraction, and handling dynamic pages—while covering essential modules, HTTP details, common obstacles like JavaScript rendering, anti‑scraping measures, and strategies for large‑scale, distributed crawling.

HTTPWeb Crawlingdata-extraction
0 likes · 14 min read
Mastering Web Crawlers: Core Principles, Architecture, and Modern Challenges
Architecture Digest
Architecture Digest
Nov 19, 2017 · Operations

Guiding Principles and Practices for High Availability and High Concurrency in Large‑Scale Systems

The article outlines core guiding principles, high‑availability strategies, and high‑concurrency techniques—such as stateless design, replica and isolation, quota control, monitoring, degradation, rollback, and scaling—to help engineers build resilient, scalable web architectures for massive traffic.

High Concurrencydistributed systemshigh availability
0 likes · 20 min read
Guiding Principles and Practices for High Availability and High Concurrency in Large‑Scale Systems
Dada Group Technology
Dada Group Technology
Nov 17, 2017 · Backend Development

Designing a High‑Availability Distributed ID Generator: From UUID to Snowflake

This article examines the requirements for globally unique IDs in distributed systems, compares classic generation schemes such as UUID, Flickr, Snowflake and TDDL, and details a customized Snowflake‑based implementation with ZooKeeper‑managed worker IDs, clock‑rollback handling, deployment optimizations, and JVM tuning to achieve high performance and reliability.

BackendSnowflakedistributed systems
0 likes · 15 min read
Designing a High‑Availability Distributed ID Generator: From UUID to Snowflake
Efficient Ops
Efficient Ops
Nov 15, 2017 · Big Data

How Tencent Built a 10 TB‑Per‑Day Full‑Link Log Monitoring Platform

This article explains how Tencent's ZhiYun full‑link log monitoring platform handles massive daily logs, overcomes challenges of diverse log formats, high throughput, fault‑tolerant design, and provides scalable storage, query, and alerting capabilities for distributed micro‑service environments.

Log Monitoringbig datadata pipeline
0 likes · 10 min read
How Tencent Built a 10 TB‑Per‑Day Full‑Link Log Monitoring Platform
Tongcheng Travel Technology Center
Tongcheng Travel Technology Center
Nov 15, 2017 · Backend Development

Design and Evolution of a Distributed Accounting System for High‑Volume Transaction Processing

This article details the background, architectural evolution, design challenges, and distributed implementation of an accounting system that automates the processing of millions of transaction records across thousands of accounts, highlighting how splitting accounts, workflows, and bills improves performance and reliability.

Accounting automationData Reconciliationbig data processing
0 likes · 10 min read
Design and Evolution of a Distributed Accounting System for High‑Volume Transaction Processing
JD Retail Technology
JD Retail Technology
Nov 14, 2017 · Operations

Design and Implementation of JD.com's Multi‑Active Distributed Architecture

This article details JD.com's multi-active distributed architecture, covering its evolution from single‑data‑center to multi‑region deployments, network design, leaf‑spine topology, data consistency mechanisms, application scheduling, monitoring, and disaster recovery strategies that enhance high availability and user experience.

Cloud InfrastructureMulti-Activedata consistency
0 likes · 11 min read
Design and Implementation of JD.com's Multi‑Active Distributed Architecture
Architecture Digest
Architecture Digest
Nov 14, 2017 · Backend Development

Architecture and Technical Practices of JD.com’s Jingmai Message Center

The article details the Jingmai Message Center’s end‑to‑end architecture, covering message ingestion via Anycall and MQ, protocol conversion, Netty‑based push system, Snowflake ID generation, Elasticsearch storage, multi‑level caching, distributed locking, and the overall design principles that enable a scalable, reliable messaging platform.

CachingNettySnowflake ID
0 likes · 9 min read
Architecture and Technical Practices of JD.com’s Jingmai Message Center
Java Backend Technology
Java Backend Technology
Nov 13, 2017 · Backend Development

Transforming Monolithic Websites to Scalable, High‑Performance Distributed Systems

Learn how early monolithic websites evolve into distributed architectures by splitting applications, services, and data, implementing load balancers, reverse proxies, caching, CDN, database sharding, and security measures, while focusing on performance, high availability, scalability, and extensibility for robust, high‑traffic sites.

distributed systemshigh availabilityperformance optimization
0 likes · 11 min read
Transforming Monolithic Websites to Scalable, High‑Performance Distributed Systems
ITPUB
ITPUB
Nov 13, 2017 · Big Data

How Real‑Time Big Data Stream Computing Powers Double 11 E‑Commerce Success

The article explains how NetEase’s real‑time big‑data stream computing platform, Sloth, handles massive, continuously generated data during China’s Double 11 shopping festival, covering use cases, architectural shifts from batch to incremental processing, technical challenges, and the role of stream‑SQL for easier development.

Real-Time ComputingSQLdistributed systems
0 likes · 16 min read
How Real‑Time Big Data Stream Computing Powers Double 11 E‑Commerce Success
ITPUB
ITPUB
Nov 13, 2017 · Big Data

How Real-Time Big Data Streaming Powers Double 11 E‑Commerce Success

The article explains how continuous data generation and real‑time stream processing enable e‑commerce platforms like NetEase Kaola to handle massive Double 11 traffic, showcasing use cases, architectural shifts from batch to incremental computing, and the technical challenges of latency, accuracy, and fault tolerance.

SQLdistributed systemse-commerce
0 likes · 15 min read
How Real-Time Big Data Streaming Powers Double 11 E‑Commerce Success
MaGe Linux Operations
MaGe Linux Operations
Nov 5, 2017 · Backend Development

Explore Alipay’s Core Backend Architecture Through Detailed Diagrams

This article presents a collection of Alipay’s system architecture diagrams—including settlement, customer service, processing, funds, and finance components—providing a reference view of the payment platform’s core backend structure, which remains largely unchanged despite data age.

Alipaydistributed systemspayment platform
0 likes · 2 min read
Explore Alipay’s Core Backend Architecture Through Detailed Diagrams
dbaplus Community
dbaplus Community
Nov 2, 2017 · Databases

What Makes Alibaba’s ApsaraCache, Codis, and Redisson Stand Out in the Redis Ecosystem?

This article summarizes key insights from the Redis track at the Cloud Xi Conference, covering Alibaba Cloud ApsaraCache's unique features, Redis Enterprise's market dominance and modules, Codis's evolution and asynchronous migration techniques, and Redisson's advanced Java client capabilities for distributed caching and locking.

ApsaraCacheCachingCodis
0 likes · 12 min read
What Makes Alibaba’s ApsaraCache, Codis, and Redisson Stand Out in the Redis Ecosystem?
21CTO
21CTO
Oct 26, 2017 · Backend Development

From Data Platform Battles to AI Dreams: A Senior Engineer’s 3‑Year Journey at Alibaba

A senior Alibaba engineer reflects on three years of building a large‑scale data platform, tackling distributed rate‑limiting challenges, leading cross‑regional projects, and pursuing AI research, while sharing personal insights on career growth, technical problem‑solving, and the value of continuous learning.

AI learningbig datacareer reflections
0 likes · 11 min read
From Data Platform Battles to AI Dreams: A Senior Engineer’s 3‑Year Journey at Alibaba
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Oct 12, 2017 · Backend Development

How Taobao Scaled Its Backend Architecture Over Time

This article outlines Taobao's learning objectives, traces the evolution of its backend architecture from V1.0 to V3.0, highlights the technical challenges faced at each stage, and explains the architectural decisions—such as modularization, service‑oriented frameworks, distributed storage, and large‑scale monitoring—that enabled massive scalability, reliability, and performance improvements.

Backendarchitecturebig data
0 likes · 6 min read
How Taobao Scaled Its Backend Architecture Over Time
Alibaba Cloud Developer
Alibaba Cloud Developer
Sep 30, 2017 · Databases

How PolarDB Redefines Cloud‑Native Relational Databases

This article traces the evolution of relational databases, explains the rise of cloud‑native computing, and details how Alibaba Cloud’s PolarDB combines storage‑compute separation, RDMA networking, shared‑disk architecture, and advanced replication techniques to deliver high‑performance, scalable, and cost‑effective database services.

Parallel RaftRDMARelational Databases
0 likes · 23 min read
How PolarDB Redefines Cloud‑Native Relational Databases
Architecture Digest
Architecture Digest
Sep 29, 2017 · Databases

Ensuring Consistency in Distributed Systems: From Local Transactions to Two‑Phase Commit and Compensation Mechanisms

This article examines various consistency solutions for distributed systems, including strong and eventual consistency, local database transactions, two‑phase commit, TCC, rollback mechanisms, local message tables, and compensation techniques, illustrating their trade‑offs and appropriate application scenarios.

ConsistencyTransactionscompensation
0 likes · 13 min read
Ensuring Consistency in Distributed Systems: From Local Transactions to Two‑Phase Commit and Compensation Mechanisms
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Sep 28, 2017 · Backend Development

Mastering Large-Scale Internet Architecture: From DNS to Distributed Caching

This article explores the principles and components of large‑scale internet architecture, covering goals such as low cost, high performance, availability and scalability, and detailing practical implementations of DNS, CDN, load balancing, web services, caching, proxies, indexing, and queueing to build robust, efficient systems.

distributed systemsload balancingscalability
0 likes · 44 min read
Mastering Large-Scale Internet Architecture: From DNS to Distributed Caching
JD Tech
JD Tech
Sep 26, 2017 · Cloud Computing

Impact of RDMA Technology on High‑Performance Data Centers and Its Adoption at JD.com

The article explains how RDMA (Remote Direct Memory Access) reduces CPU involvement, lowers latency, and increases bandwidth in data‑center networks, describes JD.com’s practical deployments across AI, big‑data, storage, and HPC workloads, and highlights industry trends toward broader RDMA adoption.

Cloud ComputingNetworkingRDMA
0 likes · 6 min read
Impact of RDMA Technology on High‑Performance Data Centers and Its Adoption at JD.com