Tagged articles

distributed systems

2274 articles · Page 16 of 23
Architects' Tech Alliance
Architects' Tech Alliance
Sep 2, 2020 · Fundamentals

Understanding Software Architecture: Concepts, Layers, Classifications, and Evolution

This article explains the fundamental concepts of software architecture, distinguishes system, subsystem, module, component, and framework, outlines architectural layers and classifications, describes the evolution from monolithic to distributed and micro‑service architectures, and discusses how to evaluate and avoid common design pitfalls.

Architecture Principlesdistributed systemsmicroservices
0 likes · 18 min read
Understanding Software Architecture: Concepts, Layers, Classifications, and Evolution
Sohu Tech Products
Sohu Tech Products
Sep 2, 2020 · Fundamentals

LegoOS: A Distributed Operating System for Disaggregated Hardware – Architecture and Design Overview

This article reviews the award‑winning 2018 OSDI paper LegoOS, describing its split‑kernel architecture, component‑based resource management across processors, memory and storage, and how it enables hardware disaggregation in data‑center clusters while addressing network latency and failure handling.

OS ArchitectureOperating Systemsdistributed systems
0 likes · 17 min read
LegoOS: A Distributed Operating System for Disaggregated Hardware – Architecture and Design Overview
Java Architect Essentials
Java Architect Essentials
Sep 1, 2020 · Backend Development

How to Ensure Data Consistency in Microservices: Patterns and Pitfalls

Microservice architectures struggle with traditional ACID transactions, so this article reviews local and distributed transaction basics, explains why 2PC/3PC are unsuitable, introduces the BASE model, and details four practical consistency patterns—reliable event, async event, business compensation, and TCC—highlighting their mechanisms, advantages, and drawbacks.

BASE theoryTCCdata consistency
0 likes · 17 min read
How to Ensure Data Consistency in Microservices: Patterns and Pitfalls
Top Architect
Top Architect
Sep 1, 2020 · Backend Development

Improving Order Number Generation to Avoid Duplicates in High‑Concurrency Java Applications

The article analyzes a real‑world incident where duplicate order IDs were generated under high concurrency, demonstrates the shortcomings of the original Java implementation, and presents a thread‑safe redesign using AtomicInteger, Java 8 date‑time API and container IP suffix to guarantee unique identifiers across parallel requests and clustered instances.

BackendJavaconcurrency
0 likes · 11 min read
Improving Order Number Generation to Avoid Duplicates in High‑Concurrency Java Applications
Top Architect
Top Architect
Sep 1, 2020 · Game Development

Why Game Companies’ Servers Are Reluctant to Adopt Microservices

The article explains, through interview excerpts, why many game studios avoid microservice architectures for their real‑time servers, highlighting latency‑sensitive communication, stateful processing, and the overhead of distributed networking that conflict with the performance demands of modern multiplayer games.

Backendarchitecturedistributed systems
0 likes · 8 min read
Why Game Companies’ Servers Are Reluctant to Adopt Microservices
Java Backend Technology
Java Backend Technology
Aug 30, 2020 · Backend Development

How to Generate Collision‑Free Order Numbers in High‑Concurrency Java Apps

The article examines a real‑world incident where duplicate order IDs appeared under high concurrency, analyzes the shortcomings of the original timestamp‑and‑random‑based scheme, and presents a revised Java implementation using thread‑safe counters, Java 8 date‑time APIs, and IP‑based suffixes to reliably produce unique order numbers even in clustered environments.

JavaUnique IDdistributed systems
0 likes · 13 min read
How to Generate Collision‑Free Order Numbers in High‑Concurrency Java Apps
Big Data Technology & Architecture
Big Data Technology & Architecture
Aug 27, 2020 · Big Data

HBase Architecture, Components, and Operations Overview

This article provides a comprehensive overview of Apache HBase’s architecture, detailing its core components such as RegionServer, HMaster, ZooKeeper, WAL, MemStore, and HFiles, and explains key processes including read/write paths, compaction, region splitting, load balancing, and recovery mechanisms.

HBaseNoSQLRecovery
0 likes · 17 min read
HBase Architecture, Components, and Operations Overview
Programmer DD
Programmer DD
Aug 23, 2020 · Backend Development

How Short URLs Work: Theory, Design, and a Java Implementation

This article explains why short URLs are used in spam SMS, outlines their benefits, describes the basic principle of mapping long URLs to short ones, and details service design choices such as storage, one‑to‑one mapping, high‑concurrency handling, distributed ID generation, and provides a Java implementation using Redis.

Redisbackend developmentdistributed systems
0 likes · 11 min read
How Short URLs Work: Theory, Design, and a Java Implementation
Full-Stack Internet Architecture
Full-Stack Internet Architecture
Aug 19, 2020 · Backend Development

Implementing Transactional Messages with Apache RocketMQ

This article explains how to use Apache RocketMQ's 2PC-based transactional messaging feature, covering the overall workflow, key concepts such as half messages and compensation, and providing complete Java/Spring Boot code examples for producers, transaction listeners, and consumers.

2PCJavaRocketMQ
0 likes · 12 min read
Implementing Transactional Messages with Apache RocketMQ
Top Architect
Top Architect
Aug 18, 2020 · Fundamentals

Fundamentals of Distributed Systems: Models, Replication, Consistency, and Core Protocols

This comprehensive article explains the core concepts of distributed systems—including node modeling, failure types, replica strategies, consistency levels, performance metrics, data distribution techniques, lease mechanisms, quorum, logging, two‑phase commit, MVCC, Paxos, and the CAP theorem—providing a solid foundation for designing robust, scalable architectures.

CAP theoremConsensusConsistency
0 likes · 53 min read
Fundamentals of Distributed Systems: Models, Replication, Consistency, and Core Protocols
Programmer DD
Programmer DD
Aug 16, 2020 · Game Development

Why Game Servers Shy Away from Microservices: Real‑Time Constraints Explained

This article examines why many game companies avoid microservice architectures, highlighting the real‑time latency, stateful memory requirements, and networking constraints that make traditional microservice patterns unsuitable for fast, multiplayer game servers.

backend architecturedistributed systemsgame server
0 likes · 8 min read
Why Game Servers Shy Away from Microservices: Real‑Time Constraints Explained
Programmer DD
Programmer DD
Aug 13, 2020 · Backend Development

How to Solve Distributed Cache Consistency Issues with Lazy Updates and Queues

This article explains the classic Cache‑Aside pattern, analyzes common cache‑database consistency problems in high‑concurrency scenarios, and presents a lazy‑update queue solution that deletes stale cache entries, routes updates through internal JVM queues, and mitigates read‑blocking and hotspot issues.

ConsistencyQueuedistributed systems
0 likes · 11 min read
How to Solve Distributed Cache Consistency Issues with Lazy Updates and Queues
DevOps
DevOps
Aug 13, 2020 · Operations

ByteDance’s Chaos Engineering Journey: Practices, Architecture, and Future Directions

This article outlines ByteDance’s adoption of chaos engineering, describing its background, industry examples, the evolution of internal fault‑injection platforms across three generations, the fault model and center design, experiment principles, and future plans for infrastructure‑level chaos and automated diagnostics.

chaos engineeringdistributed systemsfault injection
0 likes · 21 min read
ByteDance’s Chaos Engineering Journey: Practices, Architecture, and Future Directions
Big Data Technology Architecture
Big Data Technology Architecture
Aug 12, 2020 · Databases

Core Features and Architecture of ClickHouse

ClickHouse, the high‑performance columnar OLAP DBMS behind Yandex.Metrica, combines complete DBMS capabilities, column‑oriented storage with compression, vectorized execution, flexible table engines, multi‑master clustering, and extensive SQL support, offering fast online queries and scalable distributed processing for massive data workloads.

ClickHouseColumnar DatabaseOLAP
0 likes · 28 min read
Core Features and Architecture of ClickHouse
Tencent Cloud Developer
Tencent Cloud Developer
Aug 11, 2020 · Cloud Native

Tencent TDMQ: Cloud‑Native Message Queue Architecture and Implementation

Tencent’s TDMQ is a cloud‑native, Pulsar‑based message queue that separates storage and compute, offering multi‑protocol support, strong consistency, high reliability, active‑active cross‑region replication, and read‑only broker scaling to meet financial‑grade billing workloads with billions of daily transactions.

Apache PulsarFinancial Servicescloud-native
0 likes · 26 min read
Tencent TDMQ: Cloud‑Native Message Queue Architecture and Implementation
New Oriental Technology
New Oriental Technology
Aug 11, 2020 · Backend Development

Engineering Case Study of New Oriental Cloud Classroom Backend Architecture and Scaling During the Pandemic

The article details how New Oriental's Cloud Classroom backend, built with Java, Spring, MySQL, Redis, Kafka, Sentinel, and other modern technologies, scaled to support millions of users and a hundred‑fold surge in demand during the pandemic through architectural optimizations, distributed caching, traffic control, and rapid performance improvements.

JavaKafkaRedis
0 likes · 7 min read
Engineering Case Study of New Oriental Cloud Classroom Backend Architecture and Scaling During the Pandemic
Java Backend Technology
Java Backend Technology
Aug 6, 2020 · Fundamentals

Why ZooKeeper? Unveiling the Core of Distributed Coordination Services

This article explains what ZooKeeper is, why it was created, its key features such as high performance, high availability, and strong consistency, and how it simplifies distributed application development by providing coordination primitives like naming, locks, leader election, and configuration management.

ConsensusCoordination Servicedistributed systems
0 likes · 10 min read
Why ZooKeeper? Unveiling the Core of Distributed Coordination Services
Efficient Ops
Efficient Ops
Aug 3, 2020 · Backend Development

Mastering Kafka Producer API: Tips, Configurations, and Common Pitfalls

This article provides a comprehensive guide to Kafka's producer API, covering core concepts, client‑side workflow, essential configurations, idempotent and transactional producers, and practical Java code examples to help developers avoid common pitfalls and optimize message publishing.

JavaKafkaProducer API
0 likes · 21 min read
Mastering Kafka Producer API: Tips, Configurations, and Common Pitfalls
Programmer DD
Programmer DD
Aug 2, 2020 · Backend Development

How to Solve Session Sharing and Cache Issues in Distributed Systems

This article explains how to handle session sharing in distributed systems through replication, sticky sessions, or a dedicated session server, and addresses cache penetration and avalanche issues with practical mitigation techniques such as request filtering, empty-object caching, hash repositories, Redis clustering, staggered expirations, and concurrency controls.

Session Managementbackend developmentdistributed systems
0 likes · 6 min read
How to Solve Session Sharing and Cache Issues in Distributed Systems
Architects Research Society
Architects Research Society
Jul 29, 2020 · Big Data

Static Members and Incremental Cooperative Rebalancing in Apache Kafka

Apache Kafka 2.3 introduced static members and incremental cooperative rebalancing to reduce disruptive global rebalances, allowing workers to retain assignments during failures, schedule delayed rebalances, and improve scalability for Kafka Connect clusters, balancing availability and fault tolerance.

Apache KafkaIncremental RebalancingKafka Connect
0 likes · 12 min read
Static Members and Incremental Cooperative Rebalancing in Apache Kafka
Big Data Technology & Architecture
Big Data Technology & Architecture
Jul 24, 2020 · Big Data

Key Concepts and Internal Mechanisms of Apache Kafka

This article provides an in‑depth overview of Apache Kafka’s internal topics, preferred replicas, partition allocation mechanisms, log directory structure, index files, offset and timestamp lookup, log retention and compaction policies, storage architecture, delayed operations, controller role, consumer rebalance process, and producer idempotence.

Consumer RebalanceIdempotenceKafka
0 likes · 18 min read
Key Concepts and Internal Mechanisms of Apache Kafka
Amap Tech
Amap Tech
Jul 23, 2020 · Big Data

Overview of Apache Big Data Ecosystem Tools

The article surveys the Apache big‑data ecosystem, covering Hadoop’s storage and resource management, column stores HBase and Kudu, compute engines Spark, Flink, Impala, and Presto, coordination via ZooKeeper, ingestion with Sqoop and Flume, messaging Kafka, security Ranger and Sentry, metadata Atlas, OLAP Kylin, Hive, quality tool Griffin, notebooks Zeppelin, visualizations Superset and Tableau, the TPCx‑BB benchmark, and ends with an Alibaba analysis competition notice.

AnalyticsApachedata governance
0 likes · 19 min read
Overview of Apache Big Data Ecosystem Tools
转转QA
转转QA
Jul 23, 2020 · Operations

Building a Near Real‑Time Log Collection and Query System for Distributed Deployment

The article describes how a distributed deployment platform built a centralized Elasticsearch‑based log collection and query system to replace manual multi‑machine log inspection, detailing the background challenges, architecture, implementation steps, practical usage, and future improvements.

ElasticSearchKibanaLog Management
0 likes · 6 min read
Building a Near Real‑Time Log Collection and Query System for Distributed Deployment
Big Data Technology & Architecture
Big Data Technology & Architecture
Jul 22, 2020 · Big Data

Kafka Architecture and Core Concepts: Producers, Brokers, and Consumers

This article explains Kafka's fundamental architecture, including the roles of producers, brokers, and consumers, key concepts such as topics, partitions, replicas, ISR, and controller, as well as detailed mechanisms of producer client structure, interceptors, serializers, partitioners, and consumer group rebalancing strategies.

Kafkabig datadistributed systems
0 likes · 22 min read
Kafka Architecture and Core Concepts: Producers, Brokers, and Consumers
Alibaba Cloud Developer
Alibaba Cloud Developer
Jul 22, 2020 · Big Data

Exploring the Apache Big Data Ecosystem: Hadoop, Spark, Flink, and More

This article surveys the rapidly evolving big data landscape by reviewing a wide range of Apache projects—including Hadoop, Spark, Flink, HBase, Kudu, Impala, Kafka, and others—detailing their core components, architectures, strengths, and typical use‑cases for building distributed data platforms.

ApacheData ProcessingStorage
0 likes · 20 min read
Exploring the Apache Big Data Ecosystem: Hadoop, Spark, Flink, and More
dbaplus Community
dbaplus Community
Jul 19, 2020 · Databases

How Beike Achieved Millisecond Queries on a 48‑Billion‑Triple Graph with Dgraph

This article details Beike's journey of storing and querying a 480‑billion‑triple industry graph in milliseconds, covering graph database fundamentals, a comparative evaluation of JanusGraph and Dgraph, the design and deployment of a Docker‑K8s based Dgraph platform, data ingestion pipelines, a custom Graph‑SQL layer, performance testing, optimizations, and future roadmap.

BeikeDgraphGraph SQL
0 likes · 25 min read
How Beike Achieved Millisecond Queries on a 48‑Billion‑Triple Graph with Dgraph
Architect's Alchemy Furnace
Architect's Alchemy Furnace
Jul 19, 2020 · Fundamentals

Understanding Consistent Hashing: Principles, Design, and Real-World Applications

This article explains the fundamentals of hash functions, outlines the key characteristics of a good hash algorithm, and dives deep into consistent hashing—its background, mechanism, desirable properties, fault tolerance, scalability, and the use of virtual nodes to solve data skew in distributed systems.

consistent hashingdistributed systemshashing
0 likes · 12 min read
Understanding Consistent Hashing: Principles, Design, and Real-World Applications
DataFunTalk
DataFunTalk
Jul 18, 2020 · Databases

Core Features and Architecture of ClickHouse: An In‑Depth Overview

This article provides a comprehensive technical overview of ClickHouse, covering its complete DBMS capabilities, column‑oriented storage and compression, vectorized execution engine, relational SQL support, diverse table engines, multi‑master clustering, sharding, and the design philosophies that make it exceptionally fast for large‑scale analytical workloads.

ClickHouseColumnar DatabaseOLAP
0 likes · 29 min read
Core Features and Architecture of ClickHouse: An In‑Depth Overview
Sohu Tech Products
Sohu Tech Products
Jul 15, 2020 · Backend Development

Design and Implementation of a High‑Performance Global Unique ID Generation Algorithm (Mist) and Service (Medis)

This article analyses the limitations of existing global unique ID solutions such as Snowflake, presents the design of a new high‑performance, time‑independent Mist algorithm with extended lifespan, and describes its practical Redis‑backed service Medis, including performance benchmarks and architectural considerations.

RedisSnowflakedistributed systems
0 likes · 14 min read
Design and Implementation of a High‑Performance Global Unique ID Generation Algorithm (Mist) and Service (Medis)
Architects Research Society
Architects Research Society
Jul 15, 2020 · Big Data

Introduction to Apache Kafka: A Distributed Streaming Platform

This article provides a comprehensive overview of Apache Kafka, explaining its distributed, fault‑tolerant architecture, horizontal scalability, disk‑based commit log, replication mechanisms, Streams API, KSQL, and why it is widely adopted as the backbone of event‑driven, high‑throughput systems.

KafkaMessage Queuedistributed systems
0 likes · 15 min read
Introduction to Apache Kafka: A Distributed Streaming Platform
Top Architect
Top Architect
Jul 14, 2020 · Databases

Understanding Alipay’s LDC Architecture, Unitization, and CAP Analysis

The article explains how Alipay achieves massive payment throughput during Double‑11 by using logical data centers (LDC), unit‑based system design, multi‑active disaster‑recovery, and CAP‑theorem analysis, highlighting the role of OceanBase and PAXOS in ensuring consistency and availability.

CAP theoremHigh TPSLDC
0 likes · 37 min read
Understanding Alipay’s LDC Architecture, Unitization, and CAP Analysis
Selected Java Interview Questions
Selected Java Interview Questions
Jul 12, 2020 · Backend Development

Evolution of High‑Concurrency Backend Architecture: From Single‑Machine to Cloud‑Native Solutions

The article walks through Taobao's backend architecture evolution—from a single‑machine setup to distributed caching, load balancing, database sharding, microservices, containerization, and finally cloud deployment—explaining each stage's technologies, challenges, and design principles for building scalable, highly available systems.

Cachingclouddistributed systems
0 likes · 23 min read
Evolution of High‑Concurrency Backend Architecture: From Single‑Machine to Cloud‑Native Solutions
Programmer DD
Programmer DD
Jul 11, 2020 · Fundamentals

Mastering Consistent Hashing: Balance, Monotonicity, and Minimal Data Shifts

Consistent hashing, introduced by MIT in 1997, addresses hotspot issues in distributed systems by ensuring balance, monotonicity, spread, and load properties, using a ring hash space, virtual nodes, and minimal data movement when nodes are added or removed.

consistent hashingdistributed systemsload balancing
0 likes · 10 min read
Mastering Consistent Hashing: Balance, Monotonicity, and Minimal Data Shifts
AntTech
AntTech
Jul 9, 2020 · Cloud Native

Ant Group's Financial-Grade Unitized Architecture: Design, Capabilities, and Real-World Banking Cases

This article presents Ant Group’s financial‑grade unitized architecture, detailing industry‑standard distributed models, the design of RZone/GZone/CZone, its disaster‑recovery, elasticity and gray‑release capabilities, and real‑world banking case studies demonstrating cloud‑native deployment strategies in the financial sector.

Financial Technologyarchitecturecloud native
0 likes · 11 min read
Ant Group's Financial-Grade Unitized Architecture: Design, Capabilities, and Real-World Banking Cases
Java Backend Technology
Java Backend Technology
Jul 9, 2020 · Backend Development

How to Solve Distributed Cache Consistency Issues with Lazy Updates

This article explains the Cache Aside pattern, why deleting stale cache entries is often better than updating them, and presents a queue‑based lazy‑update solution that handles simple and complex consistency problems in high‑concurrency environments while outlining practical performance considerations.

BackendCache AsideConsistency
0 likes · 11 min read
How to Solve Distributed Cache Consistency Issues with Lazy Updates
Top Architect
Top Architect
Jul 8, 2020 · Fundamentals

Distributed System Characteristics and Solutions for Distributed Transaction Consistency

This article explains the key characteristics of distributed systems, introduces the CAP and BASE theories, compares strong, weak and eventual consistency models, and reviews common distributed transaction solutions such as two‑phase commit, TCC and message‑based approaches, highlighting their trade‑offs and practical considerations.

BASE theoryCAP theoremDistributed Transactions
0 likes · 13 min read
Distributed System Characteristics and Solutions for Distributed Transaction Consistency
JD Retail Technology
JD Retail Technology
Jul 1, 2020 · Backend Development

JdHotkey: A Lightweight Hotkey Detection Framework – Architecture, Workflow, and Performance

The article introduces the JdHotkey framework, a lightweight, real‑time hotkey detection solution for Redis and other services, detailing the risks of hot keys, prior mitigation methods, core design components, operational workflow, and performance benchmarks across various hardware configurations.

HotKeyRedisdistributed systems
0 likes · 16 min read
JdHotkey: A Lightweight Hotkey Detection Framework – Architecture, Workflow, and Performance
Selected Java Interview Questions
Selected Java Interview Questions
Jun 28, 2020 · Backend Development

Rate Limiting Strategies and Guava RateLimiter for High Concurrency Traffic

This article explains the concept of high traffic, compares common mitigation techniques such as caching, degradation and rate limiting, and then details four classic rate‑limiting algorithms—counter, sliding window, leaky bucket and token bucket—followed by a practical Guava RateLimiter example and a brief note on distributed scenarios.

GuavaHigh Concurrencydistributed systems
0 likes · 7 min read
Rate Limiting Strategies and Guava RateLimiter for High Concurrency Traffic
Top Architect
Top Architect
Jun 26, 2020 · Backend Development

Design and Implementation of a Transactional Message Module Using RabbitMQ and MySQL

This article explains a lightweight transactional message solution for microservices that leverages RabbitMQ, MySQL, and Spring Boot to achieve eventual consistency, detailing design principles, compensation mechanisms, table schemas, and deployment considerations for high‑throughput asynchronous processing.

MySQLRabbitMQSpring Boot
0 likes · 9 min read
Design and Implementation of a Transactional Message Module Using RabbitMQ and MySQL
Programmer DD
Programmer DD
Jun 25, 2020 · Backend Development

How to Build a Low‑Intrusive Transactional Message System with Spring Boot and RabbitMQ

This article explains a lightweight, low‑intrusive transactional message solution for microservices using Spring Boot, RabbitMQ, and MySQL, covering design principles, table schemas, transaction synchronization, compensation mechanisms, code implementation, and deployment considerations to achieve eventual consistency without sacrificing performance.

MySQLRabbitMQSpring Boot
0 likes · 24 min read
How to Build a Low‑Intrusive Transactional Message System with Spring Boot and RabbitMQ
Alibaba Cloud Developer
Alibaba Cloud Developer
Jun 24, 2020 · Backend Development

Choosing the Right Rate‑Limiting Algorithm: Simple Window, Sliding Window, Leaky Bucket, Token Bucket & Sliding Log

This article explains the purpose of flow control, compares various rate‑limiting algorithms—including simple window, sliding window, leaky bucket, token bucket, and sliding log—provides Java interface definitions and code examples, discusses their complexity, precision, smoothness, and suitability for single‑machine and distributed scenarios, and offers practical deployment tips using Sentinel, Nginx, Guava, Tair, and Redis.

JavaRedisalgorithm
0 likes · 31 min read
Choosing the Right Rate‑Limiting Algorithm: Simple Window, Sliding Window, Leaky Bucket, Token Bucket & Sliding Log
Architect
Architect
Jun 18, 2020 · Backend Development

Applying Message Queues for Decoupling in E‑commerce Architecture

The article explains why and how to use message queues to achieve low‑coupling, better performance, fault tolerance, and eventual consistency in an e‑commerce order‑processing flow, discusses common pitfalls such as message loss and duplication, and compares popular queue products like RabbitMQ, Kafka, and RocketMQ.

KafkaMessage QueueRabbitMQ
0 likes · 10 min read
Applying Message Queues for Decoupling in E‑commerce Architecture
Full-Stack Internet Architecture
Full-Stack Internet Architecture
Jun 18, 2020 · Big Data

Kafka Interview Questions: High Availability, Reliability, Consistency, Performance, and Usage Rationale

This article explains common Kafka interview questions by analyzing the system's high‑availability design, reliability mechanisms, consistency model, performance tricks such as sequential writes and zero‑copy, and the reasons for using Kafka and message queues, providing both conceptual insight and practical details.

ConsistencyKafkadistributed systems
0 likes · 12 min read
Kafka Interview Questions: High Availability, Reliability, Consistency, Performance, and Usage Rationale
Architecture Digest
Architecture Digest
Jun 16, 2020 · Backend Development

Why Use Distributed Locks? Implementation with Redis, Redisson, and Zookeeper

The article explains the inventory oversell problem in distributed e‑commerce systems, why native JVM locks fail across multiple machines, and presents practical implementations of distributed locks using Redis (including RedLock and Redisson) and Zookeeper (with Curator), comparing their advantages and drawbacks.

Java concurrencydistributed lockdistributed systems
0 likes · 17 min read
Why Use Distributed Locks? Implementation with Redis, Redisson, and Zookeeper
dbaplus Community
dbaplus Community
Jun 10, 2020 · Backend Development

Why Distributed Core Banking Systems Are the Future of Legacy Bank Architecture

This article examines the evolution of bank core systems from monolithic designs to distributed architectures, detailing the technical rationale, modular decomposition, sharding strategies, data routing, transaction processing, and performance considerations essential for modernizing banking IT infrastructure.

Transaction Processingcore bankingdata routing
0 likes · 44 min read
Why Distributed Core Banking Systems Are the Future of Legacy Bank Architecture
Top Architect
Top Architect
Jun 10, 2020 · Fundamentals

A Comprehensive Guide to Learning Distributed Systems

This article provides a thorough overview of distributed systems, explaining their definition, core challenges, key characteristics, essential components, common protocols, and practical implementations to help readers build a solid, structured learning path for mastering distributed architectures.

distributed systemsfault tolerancesystem design
0 likes · 16 min read
A Comprehensive Guide to Learning Distributed Systems
Xianyu Technology
Xianyu Technology
Jun 9, 2020 · Backend Development

Xianyu Coin Service Migration: A Service‑Based Data Migration Approach

Alibaba migrated Xianyu Coin’s points from the legacy KingTower platform to the new Banliang service using a four‑phase, service‑based approach—preparation, active/passive data migration with distributed locks, dual‑write for consistency, reconciliation, and a gradual service switch—completing the transition in a month without user impact.

Backend EngineeringRedis LockService Integration
0 likes · 11 min read
Xianyu Coin Service Migration: A Service‑Based Data Migration Approach
Architect
Architect
Jun 7, 2020 · Fundamentals

Understanding Consistency Models and Distributed Consensus Protocols

This article explains the fundamentals of distributed consistency, covering weak and strong consistency, the CAP theorem, ACID and BASE models, and detailed overviews of 2PC, 3PC, Paxos, Raft, Gossip, NWR, Quorum, and Lease mechanisms, highlighting their trade‑offs and practical use cases.

2PCCAP theoremConsensus Protocols
0 likes · 16 min read
Understanding Consistency Models and Distributed Consensus Protocols
DataFunTalk
DataFunTalk
Jun 7, 2020 · Databases

ByteKV: Design and Implementation of a Strongly Consistent Range-Partitioned KV Store

ByteKV is a C++-based, strongly consistent, range-partitioned key‑value storage system built by ByteDance, featuring a multi‑Raft consensus layer, custom storage engines (RocksDB and BlockDB), automatic partition splitting/merging, load balancing, distributed transactions, and a SQL table layer for rich data models.

ConsensusPartitioningSQL
0 likes · 47 min read
ByteKV: Design and Implementation of a Strongly Consistent Range-Partitioned KV Store
Architecture Digest
Architecture Digest
Jun 6, 2020 · Backend Development

Evolution of Project Architecture and Glossary of Common Distributed System Terms

This article explains the evolution of software project architectures—from single‑server monoliths to MVC, RPC, SOA, and micro‑services—while providing clear definitions of key terms such as clusters, load balancing, caching, and flow control for readers unfamiliar with high‑concurrency and distributed systems.

architecturedistributed systemsmicroservices
0 likes · 12 min read
Evolution of Project Architecture and Glossary of Common Distributed System Terms
Programmer DD
Programmer DD
Jun 5, 2020 · Operations

Why ZooKeeper Fails as Service Discovery: Alibaba’s 10‑Year Lessons

This article examines a decade of Alibaba’s experience with ZooKeeper‑based service discovery, arguing that ZooKeeper’s strong consistency and limited scalability make it unsuitable as a registration center and outlining design principles that favor availability, eventual consistency, and richer health‑check mechanisms.

CAP theoremdistributed systemsregistration center
0 likes · 20 min read
Why ZooKeeper Fails as Service Discovery: Alibaba’s 10‑Year Lessons
Architecture Digest
Architecture Digest
Jun 3, 2020 · Operations

Why ZooKeeper Is Not the Best Choice for Service Discovery: Design Considerations for a Registration Center

Drawing on Alibaba's decade‑long experience, this article analyses service‑discovery requirements, CAP trade‑offs, consistency versus availability, health‑check design, disaster recovery, and exception handling to argue that ZooKeeper, while excellent for coordination, is often unsuitable as the primary registration center for large‑scale microservice environments.

CAP theoremdistributed systemsregistration center
0 likes · 18 min read
Why ZooKeeper Is Not the Best Choice for Service Discovery: Design Considerations for a Registration Center
Cloud Native Technology Community
Cloud Native Technology Community
Jun 1, 2020 · Cloud Native

From Business Pain to a Fully Realized Cloud‑Native Architecture: A Step‑by‑Step Blueprint

This article walks through a practical, step‑by‑step transformation from a monolithic application to a cloud‑native, micro‑service architecture, covering planning, domain‑driven design, continuous integration, service registration, API gateways, databases, caching, logging, configuration management, containerization, performance monitoring, service governance, GitOps, traffic shading, service mesh, stress testing, and multi‑datacenter deployment.

CI/CDDevOpsIaC
0 likes · 58 min read
From Business Pain to a Fully Realized Cloud‑Native Architecture: A Step‑by‑Step Blueprint
Programmer DD
Programmer DD
May 29, 2020 · Backend Development

Demystifying Clusters, Load Balancing & Caching in Modern Backend

This article walks through the evolution of project architectures—from single‑server MVC to RPC, SOA, and micro‑services—explaining key concepts such as clusters, load‑balancing strategies, and various caching mechanisms, helping readers grasp how high‑concurrency, distributed systems are designed and optimized.

Cachingdistributed systemsload balancing
0 likes · 12 min read
Demystifying Clusters, Load Balancing & Caching in Modern Backend
dbaplus Community
dbaplus Community
May 25, 2020 · Operations

Scaling CAT Monitoring at Ctrip: Thread Model, Client Computation & Memory Tweaks

This article details how Ctrip optimized the CAT monitoring system—covering its large‑scale deployment, thread‑model redesign, offloading calculations to clients, double‑buffered reporting, and string handling improvements—to dramatically cut CPU usage, GC pressure, and memory consumption while handling billions of messages daily.

GCJavadistributed systems
0 likes · 25 min read
Scaling CAT Monitoring at Ctrip: Thread Model, Client Computation & Memory Tweaks
Yanxuan Tech Team
Yanxuan Tech Team
May 25, 2020 · Operations

How NetEase Cloud Music Built a Scalable Full‑Link Tracing System for Real‑Time Service Diagnosis

This article details the design, implementation, and evolution of NetEase Cloud Music's full‑link tracing platform, covering its motivations, architecture, low‑overhead data collection, multi‑dimensional analysis, service grooming, automated diagnosis, and future plans for AI‑driven anomaly detection and big‑data processing.

distributed systemsobservabilityservice monitoring
0 likes · 19 min read
How NetEase Cloud Music Built a Scalable Full‑Link Tracing System for Real‑Time Service Diagnosis
macrozheng
macrozheng
May 21, 2020 · Big Data

Mastering Kafka: Core Concepts, Architecture, and Reliability Guarantees

This comprehensive guide covers Kafka's definition, publish/subscribe model, key components, storage mechanisms, producer and consumer strategies, and reliability features such as ACK levels, ISR, and exactly‑once semantics, providing a solid foundation for real‑time big‑data processing.

KafkaMessage Queuebig data
0 likes · 16 min read
Mastering Kafka: Core Concepts, Architecture, and Reliability Guarantees
Youzan Coder
Youzan Coder
May 20, 2020 · Backend Development

Real-Time Loss Prevention System: Architecture and Implementation at YouZan

YouZan’s real‑time loss‑prevention platform monitors database binlogs, transforms and verifies transaction data across five loosely coupled layers, handling 200 million daily messages and 60 million checks with dynamic sharding, caching and distributed locks to detect over‑charges, duplicate refunds, migration inconsistencies and unauthorized data changes.

HBaseMessage QueueSharding Strategy
0 likes · 12 min read
Real-Time Loss Prevention System: Architecture and Implementation at YouZan
Xiaokun's Architecture Exploration Notes
Xiaokun's Architecture Exploration Notes
May 15, 2020 · Backend Development

Mastering Distributed System Design: Key Principles, Techniques, and Best Practices

This comprehensive guide explains why distributed systems are needed, outlines design goals, explores essential technologies and architectural patterns, and provides practical strategies for scalability, high availability, service governance, DevOps automation, and monitoring to help engineers build robust distributed architectures.

Service Governancedistributed systemshigh availability
0 likes · 22 min read
Mastering Distributed System Design: Key Principles, Techniques, and Best Practices
Meituan Technology Team
Meituan Technology Team
May 14, 2020 · Cloud Native

Meituan Naming Service (MNS) 2.0: Architecture Evolution and Business Enablement

Meituan’s Naming Service 2.0 replaces the ZooKeeper‑based 1.0 design with a four‑layer, AP‑oriented architecture that leverages a service‑mesh sidecar, sharded KV storage, and a control service layer, delivering eight‑fold throughput gains, sub‑second latency, zero‑downtime migration for most services, and new business capabilities such as traffic isolation, elastic scaling, and data‑driven SLA monitoring.

Service Meshcloud nativedistributed systems
0 likes · 25 min read
Meituan Naming Service (MNS) 2.0: Architecture Evolution and Business Enablement
Tencent Tech
Tencent Tech
May 11, 2020 · Big Data

How Tencent Scaled Elasticsearch to Thousands of Nodes: Core Kernel Optimizations Revealed

This article details Tencent's large‑scale Elasticsearch deployment, covering its massive usage scenarios, the availability, performance, cost and scalability challenges faced, and the comprehensive kernel‑level optimizations—including memory‑based throttling, storage‑model merging, off‑heap caching, rollup and metadata improvements—that enable PB‑level clusters with high reliability and low expense.

ElasticSearchbig datadistributed systems
0 likes · 27 min read
How Tencent Scaled Elasticsearch to Thousands of Nodes: Core Kernel Optimizations Revealed
Architecture Digest
Architecture Digest
May 11, 2020 · Backend Development

Ensuring Idempotency in Distributed Systems: Unique ID Generation Strategies

The article explains why idempotency is essential for reliable service calls, discusses using unique identifiers such as UUIDs and Snowflake algorithms, compares centralized and client‑side ID generation, and offers practical storage and query‑optimisation techniques to prevent duplicate orders and resource waste.

IdempotencySnowflakeUUID
0 likes · 6 min read
Ensuring Idempotency in Distributed Systems: Unique ID Generation Strategies
Java Captain
Java Captain
May 8, 2020 · Big Data

Elasticsearch Adoption and Architecture Cases in Major Chinese Companies

The article surveys how leading Chinese tech firms such as JD Daojia, Ctrip, Qunar, 58.com, and Didi have adopted Elasticsearch for large‑scale search, real‑time analytics, and security, detailing their evolving cluster architectures, shard strategies, data volumes, and supporting services.

ElasticSearcharchitecturebig data
0 likes · 11 min read
Elasticsearch Adoption and Architecture Cases in Major Chinese Companies
Programmer DD
Programmer DD
May 8, 2020 · Backend Development

How to Become a Middleware Engineer: Skills, Roadmap, and Tips

This article outlines what middleware development entails, the essential technical and professional qualities required, various types of middleware, and a step‑by‑step learning roadmap for Java developers aiming to break into middleware engineering.

Career GuideDevOpsJava
0 likes · 8 min read
How to Become a Middleware Engineer: Skills, Roadmap, and Tips
Tencent Cloud Developer
Tencent Cloud Developer
Apr 29, 2020 · Cloud Computing

Large-Scale Task Scheduling Architecture of Tencent Meeting and VStation

The talk explains how Tencent’s self‑developed VStation scheduler, integrated with TKE and using a hybrid sharding‑plus‑master‑worker architecture, enabled Tencent Meeting to scale to over 100 000 hosts and one million CPU cores, cutting provisioning time to under ten seconds while handling thousands of tasks per minute through DAG‑driven automation and fault‑tolerant mechanisms.

Tencent MeetingVStationdistributed systems
0 likes · 24 min read
Large-Scale Task Scheduling Architecture of Tencent Meeting and VStation
Programmer DD
Programmer DD
Apr 29, 2020 · Operations

How to Keep Your Distributed System Running Even When Upstream Services Fail

The article explains why distributed systems must stay alive despite upstream or downstream failures, emphasizing rate limiting and circuit breaking as essential practices to prevent fault propagation and ensure service reliability, and it invites developers to assess their own safeguards.

Circuit Breakingdistributed systemsrate limiting
0 likes · 3 min read
How to Keep Your Distributed System Running Even When Upstream Services Fail
Tencent Cloud Developer
Tencent Cloud Developer
Apr 28, 2020 · Big Data

Evolution of Ctrip Vacation Pricing Engine: Architecture, Challenges, and Optimizations

Ctrip’s vacation pricing engine evolved from a MySQL‑based synchronous queue to a Kafka‑driven, Spark‑parallelized architecture using HBase, dramatically cutting task generation from five hours to 1.5 hours, boosting price‑accuracy above 90 % while handling billions of calculations and external API constraints.

KafkaSparkdistributed systems
0 likes · 18 min read
Evolution of Ctrip Vacation Pricing Engine: Architecture, Challenges, and Optimizations
Top Architect
Top Architect
Apr 28, 2020 · Operations

Evolution of System Architecture: From Single‑Machine to Distributed Designs

The article outlines the historical evolution of IT architecture—from early single‑machine deployments through hot‑standby and multi‑node active clusters to modern distributed systems—explaining the motivations, trade‑offs, and key technologies that drive each generation and offering guidance on selecting the right architecture for business needs.

databasedistributed systemssystem design
0 likes · 8 min read
Evolution of System Architecture: From Single‑Machine to Distributed Designs
Tencent Cloud Developer
Tencent Cloud Developer
Apr 26, 2020 · Backend Development

Design and Evolution of Ctrip Flight Search System: High‑Throughput Caching, Real‑Time Computing, Load Balancing and AI

Ctrip’s flight search service processes two billion daily queries by employing a multi‑level Redis cache, machine‑learning‑driven TTLs, distributed pooling and overload protection, AI‑based anti‑scraping, and robust load‑balancing across three data centers, delivering sub‑second latency, up to three‑fold throughput gains and significant cost reductions.

AICachingReal-Time Computing
0 likes · 23 min read
Design and Evolution of Ctrip Flight Search System: High‑Throughput Caching, Real‑Time Computing, Load Balancing and AI
Java Backend Technology
Java Backend Technology
Apr 26, 2020 · Databases

When to Shard Your Database? A Practical Guide to Partitioning Strategies

This article explains database bottlenecks caused by IO and CPU limits, introduces horizontal and vertical sharding for databases and tables, compares popular sharding tools, discusses challenges such as distributed transactions, cross‑node joins, pagination and global ID generation, and offers guidance on when and how to apply sharding in real‑world systems.

Partitioningdatabasedistributed systems
0 likes · 14 min read
When to Shard Your Database? A Practical Guide to Partitioning Strategies
Top Architect
Top Architect
Apr 23, 2020 · Backend Development

Evolution and Core Principles of Large-Scale Website Architecture

The article outlines how large‑scale websites evolve from single‑server monoliths to layered, distributed architectures by separating application, service and data layers, adding caching, clustering, load balancing, CDNs, NoSQL, micro‑services and automation to achieve performance, high availability, scalability, extensibility and security.

distributed systemswebsite architecture
0 likes · 17 min read
Evolution and Core Principles of Large-Scale Website Architecture
360 Tech Engineering
360 Tech Engineering
Apr 21, 2020 · Backend Development

Using ETCD for Leader Election and High Availability: Architecture, Installation, and Go Implementation

This article explains ETCD's role as a distributed key‑value store, details its architecture and leader election mechanism, provides step‑by‑step cluster deployment on CentOS with systemd, and demonstrates a Go implementation of leader election to achieve high availability.

Leader Electiondistributed systemshigh availability
0 likes · 10 min read
Using ETCD for Leader Election and High Availability: Architecture, Installation, and Go Implementation
ITFLY8 Architecture Home
ITFLY8 Architecture Home
Apr 21, 2020 · Backend Development

8 Design Principles for Business Middle Platforms and Distributed Services

This article explains how to abstract business functions into middle‑platform models, outlines eight essential design principles for services, and describes key distributed mechanisms such as service registration, elastic scaling, rate limiting, gray‑release, message queues, and transaction handling to build robust, scalable backend systems.

Backenddistributed systemsmicroservices
0 likes · 18 min read
8 Design Principles for Business Middle Platforms and Distributed Services
Java Backend Technology
Java Backend Technology
Apr 14, 2020 · Cloud Native

Why ZooKeeper Isn’t the Best Choice for Service Discovery: Design Insights

This article analyzes the limitations of ZooKeeper for service discovery, covering consistency, partition tolerance, scalability, persistence, health‑checking, disaster‑recovery, and operational complexities, and explains why modern registration centers should favor AP designs and richer health‑check mechanisms.

CAP theoremZooKeeperdistributed systems
0 likes · 19 min read
Why ZooKeeper Isn’t the Best Choice for Service Discovery: Design Insights
Tencent Cloud Developer
Tencent Cloud Developer
Apr 10, 2020 · Backend Development

High‑Availability Architecture for a Flash‑Sale System in the Weishi Spring Festival Card‑Collect Event

The article details a high‑availability flash‑sale architecture for Weishi’s Spring Festival card‑collect event, describing three design models, a funnel‑style traffic‑filtering approach, product and client strategies, layered rate limiting, sharding, asynchronous order handling, multi‑region DB redundancy, and a three‑level degradation plan to sustain extreme concurrency.

distributed systemsflash salehigh availability
0 likes · 9 min read
High‑Availability Architecture for a Flash‑Sale System in the Weishi Spring Festival Card‑Collect Event
Tencent Cloud Developer
Tencent Cloud Developer
Apr 9, 2020 · Cloud Computing

Twenty-Seven New Tencent Cloud TVP Experts Join to Advance Cloud Computing

The article announces that twenty‑seven distinguished professionals from fields such as blockchain, AI, big data, databases, and distributed systems have joined Tencent Cloud’s Valuable Professional program, highlighting their backgrounds, achievements, and commitment to advancing cloud computing and enriching the developer ecosystem.

TVPTencent Cloudblockchain
0 likes · 19 min read
Twenty-Seven New Tencent Cloud TVP Experts Join to Advance Cloud Computing
Top Architect
Top Architect
Apr 9, 2020 · Backend Development

Low‑Latency and High‑Availability Design of RocketMQ: Evolution, Optimizations, and Capacity Planning

This article reviews the evolution of Alibaba's Aliware message engine, analyzes the low‑latency and high‑availability challenges faced during Double 11, and details the architectural, JVM, memory, rate‑limiting, and multi‑replica solutions that enabled RocketMQ to achieve sub‑millisecond write latency and five‑nine availability.

RocketMQcapacity planningdistributed systems
0 likes · 29 min read
Low‑Latency and High‑Availability Design of RocketMQ: Evolution, Optimizations, and Capacity Planning
21CTO
21CTO
Apr 6, 2020 · Operations

How Alipay Achieved Near‑Zero Downtime with Multi‑Datacenter Failover Architecture

This article explains the evolution of Alipay's high‑availability and disaster‑recovery architecture—from a simple single‑datacenter design to a multi‑datacenter, unit‑based system with failover and blue‑green deployment—highlighting the challenges, solutions, and operational benefits that enable continuous service during massive traffic spikes.

Alipay architectureBlue-Green DeploymentCloud Operations
0 likes · 17 min read
How Alipay Achieved Near‑Zero Downtime with Multi‑Datacenter Failover Architecture
Wukong Talks Architecture
Wukong Talks Architecture
Apr 6, 2020 · Fundamentals

Fundamentals of Distributed Systems: Microservices, Clustering, Load Balancing, Service Registry, Configuration Center, Circuit Breaker, and API Gateway

This article introduces core concepts of distributed systems, covering microservices, clustering, remote invocation, load balancing algorithms, service registration and discovery, configuration management, circuit breaking, degradation strategies, and API gateways, providing a comprehensive overview for building resilient cloud-native applications.

circuit breakerdistributed systemsload balancing
0 likes · 6 min read
Fundamentals of Distributed Systems: Microservices, Clustering, Load Balancing, Service Registry, Configuration Center, Circuit Breaker, and API Gateway
Continuous Delivery 2.0
Continuous Delivery 2.0
Apr 3, 2020 · Operations

Scalable and Reliable Configuration Distribution at Facebook

This article explains how Facebook’s Configerator system achieves scalable, reliable configuration distribution using a push model, a hierarchical Zeus tree, Package Vessel for large data, and multi‑repo Git strategies to improve commit throughput and fault tolerance.

configuration managementdistributed systemspush model
0 likes · 11 min read
Scalable and Reliable Configuration Distribution at Facebook
Programmer DD
Programmer DD
Mar 28, 2020 · Backend Development

Why Is Kafka So Fast? Uncover the 11 Performance Secrets

Kafka achieves its remarkable speed by combining sequential I/O, batch processing, compression, zero‑copy, careful client‑side work, and a design that avoids costly fsync and garbage collection, while maintaining durability, ordering, and at‑least‑once delivery, making it a high‑throughput, low‑latency event streaming platform.

Batch ProcessingKafkaMessage Queue
0 likes · 15 min read
Why Is Kafka So Fast? Uncover the 11 Performance Secrets