Tagged articles

high availability

1528 articles · Page 1 of 16
Random Bulletin
Random Bulletin
Oct 4, 2026 · Backend Development

From Hours to Minutes: Pre-Positioning Decisions to Slash Recovery Time at 10M QPS

This article decomposes recovery time into detection, decision, execution, and verification phases, showing how 10M QPS systems reduce recovery from hours to minutes by pre-positioning decisions at design time, bounding data loss via semi-sync replication, automating failover loops with fencing and quorum, and validating RTO through disciplined drills.

RTOcapacity-planningchaos engineering
0 likes · 34 min read
From Hours to Minutes: Pre-Positioning Decisions to Slash Recovery Time at 10M QPS
Random Bulletin
Random Bulletin
Oct 1, 2026 · Backend Development

Fault Domain Design: Turning Blast Radius from 100% into a Tunable 1/N Parameter

The article presents a layered fault-domain strategy—physical anti-affinity, logical isolation (sharding, cluster groups, swimlanes, bulkheads), cell-based architecture, chaos-engineering validation, and quantitative governance metrics—to shrink the blast radius of a ten-million-QPS system from a fixed 100% to a controllable 1/N design parameter.

KubernetesSLOanti-affinity
0 likes · 25 min read
Fault Domain Design: Turning Blast Radius from 100% into a Tunable 1/N Parameter
liandk
liandk
Sep 29, 2026 · Backend Development

Java Fault Analysis Masterclass: 7-Step Review, 6 Root Causes, 8 HA Principles

This final chapter of a 20-part Java series presents a 7-step fault review SOP, categorizes 99% of production faults into six root causes, defines eight high-availability architecture principles, and outlines a four-layer risk prevention system to shift from reactive firefighting to proactive stability.

Architecture PrinciplesFault AnalysisJava
0 likes · 13 min read
Java Fault Analysis Masterclass: 7-Step Review, 6 Root Causes, 8 HA Principles
Linux Tech Enthusiast
Linux Tech Enthusiast
Sep 27, 2026 · Operations

LVS vs Nginx: Layer 4 vs Layer 7 Load Balancing Trade-offs

This article compares LVS (Layer 4) and Nginx (Layer 7) load balancers, explaining why LVS achieves higher throughput by only inspecting IP headers while Nginx terminates TCP connections for HTTP-aware routing, health checks, and request retry — at the cost of added latency and configuration complexity.

LVSNginxhealth checks
0 likes · 11 min read
LVS vs Nginx: Layer 4 vs Layer 7 Load Balancing Trade-offs
Random Bulletin
Random Bulletin
Sep 18, 2026 · Operations

Fault Drills at 10M QPS: From Zero to Regular Cadence

This article explains how to evolve fault drills from one-off exercises into a regular engineering practice, covering risk mapping, safety boundaries, Game Day execution, metrics, scenario libraries, and integrating remediation into daily workflows for high-QPS systems.

Game DaySREchaos engineering
0 likes · 31 min read
Fault Drills at 10M QPS: From Zero to Regular Cadence
liandk
liandk
Sep 15, 2026 · Databases

MySQL MGR Cluster: Paxos Consensus, Zero-Downtime HA & Brain-Split Proof Architecture

This article explains MySQL Group Replication (MGR), a native high-availability cluster based on Paxos consensus, covering its architecture, single-primary vs multi-primary modes, automatic failover, brain-split prevention via majority voting, production deployment rules, and common pitfalls to avoid for zero-data-loss transactional systems.

Brain SplitDatabase ClusteringGroup Replication
0 likes · 12 min read
MySQL MGR Cluster: Paxos Consensus, Zero-Downtime HA & Brain-Split Proof Architecture
Code Farming
Code Farming
Sep 7, 2026 · Backend Development

How Encyclopedia Systems Survive Data Center Fires: Architecture Deep Dive

This article breaks down the four-step architecture design of a high-concurrency encyclopedia system, covering latency-driven multi-data-center deployment, a five-layer request chain with caching at each level, a three-tier optimization pyramid, and master-slave synchronization with automatic failover for disaster recovery.

CDN cachingCanalGeoDNS
0 likes · 8 min read
How Encyclopedia Systems Survive Data Center Fires: Architecture Deep Dive
Fei's Miscellaneous Talks
Fei's Miscellaneous Talks
Sep 6, 2026 · Artificial Intelligence

How Prediction Servers Score 1000 Candidates in Milliseconds: Fine-Ranking Architecture Deep Dive

This article details the architecture and optimization of a Prediction Server for fine-ranking in recommender systems, covering model management, batch inference, multi-objective fusion, probability calibration, deployment pipelines, and high-availability patterns to score thousands of candidates within milliseconds.

ONNX Runtimebatch inferencefine ranking
0 likes · 46 min read
How Prediction Servers Score 1000 Candidates in Milliseconds: Fine-Ranking Architecture Deep Dive
liandk
liandk
Sep 3, 2026 · Databases

MySQL Read-Write Separation Masterclass: Replication, Lag Solutions & HA Architecture

This comprehensive guide covers MySQL read-write separation from fundamentals to production implementation, including master-slave replication principles, replication lag causes and three enterprise-grade solutions, one-master-multi-slave architecture patterns, MHA automatic failover, Sharding-JDBC middleware configuration, and five critical production pitfalls to avoid.

MHAMaster-Slave ReplicationMySQL
0 likes · 10 min read
MySQL Read-Write Separation Masterclass: Replication, Lag Solutions & HA Architecture
Xiaolin Talks Programming
Xiaolin Talks Programming
Aug 26, 2026 · Backend Development

Spring Boot High Availability: Nacos Service Discovery, OpenFeign & Sentinel Resilience

This article details how to build highly available Spring Boot microservices using Nacos for service discovery and configuration, OpenFeign for declarative REST calls with load balancing, timeouts, and retries, and Sentinel for circuit breaking and fallback handling, including multi-environment isolation via namespaces and groups.

Circuit BreakerNacosOpenFeign
0 likes · 20 min read
Spring Boot High Availability: Nacos Service Discovery, OpenFeign & Sentinel Resilience
Code Farming
Code Farming
Aug 21, 2026 · Databases

How to Cut Redis Cluster Failover to Under 10 Seconds

The article breaks down Redis‑Cluster failover into detection, election, failover, and client perception stages, explains the timing bottlenecks of each, and provides concrete server‑side and Lettuce client configurations that shrink end‑to‑end recovery to under ten seconds.

LettuceRedisRedis Cluster
0 likes · 7 min read
How to Cut Redis Cluster Failover to Under 10 Seconds
ITPUB
ITPUB
Aug 17, 2026 · Databases

How a Veteran DBA Tackles Financial Multi‑Database and Big Data Architecture

In this interview, senior DBA Yao Wei shares over a decade of hands‑on experience designing, testing, and operating heterogeneous database ecosystems for retail, internet, and financial services, detailing pre‑emptive DBA involvement, compatibility pitfalls, high‑availability strategies, automated validation pipelines, and migration to domestic databases.

CDCautomationdata migration
0 likes · 32 min read
How a Veteran DBA Tackles Financial Multi‑Database and Big Data Architecture
CodeSmart Hoops
CodeSmart Hoops
Aug 16, 2026 · Interview Experience

Elasticsearch Interview Self-Test: 8 ELK Questions & Answers Explained

This guide provides eight comprehensive Elasticsearch interview questions covering inverted indexes, shard sizing, ILM lifecycle, dynamic templates, processing pipelines, query DSL, scaling strategies, and high‑availability deployment, each accompanied by detailed answers, code examples, and best‑practice recommendations for ELK stack professionals.

ELKElasticsearchFilebeat
0 likes · 25 min read
Elasticsearch Interview Self-Test: 8 ELK Questions & Answers Explained
Architect's Guide
Architect's Guide
Aug 14, 2026 · Backend Development

Master Nginx in One Hour: A Quick Guide

This article introduces Nginx’s core concepts, walks through installing required packages, configuring the server, and explains key directives such as worker processes, events, and http blocks, then demonstrates practical setups for reverse proxy, load balancing, static‑dynamic separation, performance tuning, and high‑availability clustering.

ConfigurationNginxhigh availability
0 likes · 13 min read
Master Nginx in One Hour: A Quick Guide
Cloud Architecture
Cloud Architecture
Aug 12, 2026 · Databases

How Redis Cluster’s Decentralized Design Powers Billion‑Scale Traffic

When a single Redis instance can no longer hold the data volume or write load of e‑commerce workloads, the traditional master‑slave with Sentinel model reaches its limits, and Redis Cluster—by sharding data across 16,384 slots, using gossip‑based topology, and removing a central control plane—delivers horizontal scaling and fault‑tolerance for billions of requests, provided key design, hash tags, hot‑key mitigation, and client routing are applied.

Rediscachecluster
0 likes · 32 min read
How Redis Cluster’s Decentralized Design Powers Billion‑Scale Traffic
Cloud Architecture
Cloud Architecture
Aug 11, 2026 · Databases

Redis Sentinel Deep Dive: Leader Election, Failover Mechanics, and Production Best Practices

This article dissects Redis Sentinel’s high‑availability workflow—from failure detection, SDOWN/ODOWN states, and quorum logic to leader election, replica promotion, and configuration propagation—while illustrating each step with a real‑world e‑commerce cache case, detailed configuration snippets, Kubernetes deployment patterns, Spring Boot integration, and operational playbooks for observability and fault‑injection testing.

KubernetesRedisSentinel
0 likes · 48 min read
Redis Sentinel Deep Dive: Leader Election, Failover Mechanics, and Production Best Practices
Cloud Architecture
Cloud Architecture
Aug 10, 2026 · Databases

Master‑Slave Replication in Redis: Core Mechanics Explained and Production Deployment

This article provides a comprehensive, production‑focused analysis of Redis master‑slave replication, covering its internal state machine, full and partial sync processes, configuration pitfalls, performance bottlenecks, consistency trade‑offs, and practical deployment patterns with Docker, Kubernetes, and Spring Boot.

Docker ComposeKubernetesRedis
0 likes · 38 min read
Master‑Slave Replication in Redis: Core Mechanics Explained and Production Deployment
21CTO
21CTO
Aug 10, 2026 · Databases

Build a High‑Availability PostgreSQL Cluster with Streaming Replication and Read‑Write Splitting from Scratch

This guide walks through the complete process of designing, configuring, and operating a PostgreSQL primary‑standby setup with streaming replication, choosing between asynchronous and synchronous modes, implementing read‑write splitting via application routing or PgBouncer, and handling monitoring, failover, and common pitfalls for production‑grade high availability.

PatroniPostgreSQLRead-Write Splitting
0 likes · 15 min read
Build a High‑Availability PostgreSQL Cluster with Streaming Replication and Read‑Write Splitting from Scratch
Cloud Architecture
Cloud Architecture
Aug 9, 2026 · Databases

Redis Persistence Deep Dive: From a Major P0 Outage to RDB+AOF Hybrid Implementation

The article analyses a real‑world P0 outage caused by treating Redis as a simple cache, explains why persistence is the decisive factor when Redis stores session, inventory or lock data, and provides a step‑by‑step guide to RDB, AOF and hybrid persistence, configuration, monitoring, recovery and best‑practice recommendations.

AOFHybrid PersistencePerformance
0 likes · 33 min read
Redis Persistence Deep Dive: From a Major P0 Outage to RDB+AOF Hybrid Implementation
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Aug 9, 2026 · Cloud Native

How ACK One Fleet Transforms Agent Sandbox from Single-Cluster to Multi-Cluster

The article explains how ACK One Fleet upgrades the AI Agent Sandbox from a single‑cluster Kubernetes setup to a multi‑cluster architecture, addressing capacity limits, fault‑domain risks, and scheduling inefficiencies while providing global capacity control, water‑level balancing, fault‑tolerant failover, and faster sandbox startup through E2B and CRD integrations.

ACK OneE2BKubernetes
0 likes · 11 min read
How ACK One Fleet Transforms Agent Sandbox from Single-Cluster to Multi-Cluster
YiSu Grain
YiSu Grain
Aug 8, 2026 · Databases

Day 51: Database Architecture – Master‑Slave Replication, Read‑Write Separation, Sharding & Consistency

The article walks through diagnosing database bottlenecks, explains MySQL replication flow, read‑write separation benefits and limits, shows when to apply vertical versus horizontal partitioning, details sharding‑key selection, cross‑shard challenges, high‑availability steps, and presents a complete e‑commerce case study.

Database ReplicationHorizontal PartitioningRead-Write Separation
0 likes · 29 min read
Day 51: Database Architecture – Master‑Slave Replication, Read‑Write Separation, Sharding & Consistency
Cloud Architecture
Cloud Architecture
Aug 7, 2026 · Backend Development

Designing an Industrial‑Grade Message Queue for Tens of Millions of Orders

This article presents a step‑by‑step design of HermesMQ, an industrial‑grade message queue built from scratch to support ten‑million‑order traffic, covering storage as sequential logs, network architecture with Netty and Reactor, high‑availability replication, partition ordering, transaction messaging, back‑pressure, observability, and practical deployment guidelines.

JavaTransaction Messagingdistributed systems
0 likes · 45 min read
Designing an Industrial‑Grade Message Queue for Tens of Millions of Orders
Ray's Galactic Tech
Ray's Galactic Tech
Aug 7, 2026 · Operations

Scaling Nginx to Handle 500M Daily Requests: From Reverse Proxy to Traffic Governance Hub

This article walks through an enterprise‑grade Nginx upgrade for a payment platform handling ~500 million daily requests, detailing why simple reverse‑proxying fails, how Nginx can become a traffic‑governance edge with rate limiting, edge caching, gray releases, high‑availability, and observability, and provides production‑ready configurations and step‑by‑step analysis.

Edge CachingGray ReleaseNginx
0 likes · 42 min read
Scaling Nginx to Handle 500M Daily Requests: From Reverse Proxy to Traffic Governance Hub
ITPUB
ITPUB
Aug 7, 2026 · Databases

The Evolution and Practice of Multi-Write Multi-Read Shared Storage Clusters – UXDB SRAC Whitepaper

The whitepaper released by Youxuan Database outlines the market size of the DBMS industry, traces the technical evolution of multi‑write multi‑read shared‑storage clusters from early research to modern implementations, details the UXDB SRAC architecture and its innovations, and validates its feasibility through certifications, benchmark results, and real‑world case studies.

UXDB SRACcache fusioncluster database
0 likes · 9 min read
The Evolution and Practice of Multi-Write Multi-Read Shared Storage Clusters – UXDB SRAC Whitepaper
liandk
liandk
Aug 5, 2026 · Databases

Mastering Redis High Availability: Replication, Sentinel, and Cluster Explained

The article explains why Redis must be highly available and walks through three progressive architectures—master‑slave replication, Sentinel automatic failover, and Redis Cluster—detailing their mechanisms, advantages, drawbacks, and when to choose each for small, medium, or large‑scale production systems.

Database ScalingRedisReplication
0 likes · 7 min read
Mastering Redis High Availability: Replication, Sentinel, and Cluster Explained
Golang Shines
Golang Shines
Aug 3, 2026 · Cloud Native

How I Built a Production‑Ready HA Kubernetes Cluster in Minutes

When my manager suddenly demanded a production‑grade, highly available Kubernetes cluster integrated with a private Harbor registry, I followed a comprehensive step‑by‑step guide to finish the entire setup within a few hours, and now share the 83‑page manual for anyone to replicate.

Cluster DeploymentHarborKubernetes
0 likes · 3 min read
How I Built a Production‑Ready HA Kubernetes Cluster in Minutes
Xiaolin Talks Programming
Xiaolin Talks Programming
Aug 2, 2026 · Backend Development

Spring Boot + Spring Session: Distributed Session Management & Multi-Client Sync in Practice

This article provides a production-ready guide to replacing traditional HttpSession with Spring Session backed by Redis, covering multi-client session unification, performance tuning, security hardening, high-availability patterns, and practical Spring Boot 3.x configuration with code examples.

Distributed SessionMulti-clientRedis
0 likes · 19 min read
Spring Boot + Spring Session: Distributed Session Management & Multi-Client Sync in Practice
Ray's Galactic Tech
Ray's Galactic Tech
Aug 1, 2026 · Databases

Beyond CRUD: Full‑Scale Production Guide for MySQL 8.4 LTS

This article walks through a complete production‑grade view of MySQL 8.4 LTS, explaining how a chain of traffic spikes, connection‑pool exhaustion, long transactions and replication lag can cause an avalanche, and then detailing the five core modules, seven production mechanisms, architectural evolution steps, incident post‑mortems, and concrete configuration and code examples to build a resilient MySQL service.

InnoDBMySQLPerformance
0 likes · 36 min read
Beyond CRUD: Full‑Scale Production Guide for MySQL 8.4 LTS
CTO Full-Stack Academy
CTO Full-Stack Academy
Jul 30, 2026 · Operations

Common Cluster Issues and Practical Solutions for Apps, DBs, Caches, MQ, Files, and Search

The article enumerates typical problems encountered in application, database, cache, message‑queue, file‑server, and search clusters—such as session loss, uneven load, data inconsistency, and node failures—and provides concrete mitigation strategies like JWT authentication, distributed locks, health checks, NTP sync, and proper sharding.

Cachingclusterdatabase
0 likes · 58 min read
Common Cluster Issues and Practical Solutions for Apps, DBs, Caches, MQ, Files, and Search
Alibaba Cloud Native
Alibaba Cloud Native
Jul 30, 2026 · Cloud Native

Replication‑Free Failover for RocketMQ: Achieving Second‑Level Takeover Without Data Copy

The ACM FSE‑2026 industry paper introduces a replication‑free failover mechanism for cloud‑native stateful services like Apache RocketMQ, using protocol‑level write isolation and multi‑attach storage to achieve second‑level recovery without extra data copies, while maintaining low cost and near‑native throughput.

Protocol FencingReplication-Free FailoverRocketMQ
0 likes · 8 min read
Replication‑Free Failover for RocketMQ: Achieving Second‑Level Takeover Without Data Copy
YiSu Grain
YiSu Grain
Jul 27, 2026 · Fundamentals

Why 10 Gbps Hospital Networks Still Lag – OSI, TCP/UDP, QoS & High‑Availability Explained

Even with a 10 Gbps backbone, hospital networks can still suffer latency, jitter, packet loss and outages; this article walks through why layering (OSI/TCP‑IP), choosing TCP or UDP, applying QoS metrics, designing a three‑tier campus network and implementing five‑layer high‑availability to meet reliability, latency and continuity requirements.

Network ArchitectureOSIQoS
0 likes · 23 min read
Why 10 Gbps Hospital Networks Still Lag – OSI, TCP/UDP, QoS & High‑Availability Explained
Dabaoshi
Dabaoshi
Jul 26, 2026 · Databases

When to Scale MySQL: From Single Instance to Distributed Architecture

The article explains why a single MySQL server eventually hits read, write, or availability limits, outlines a step‑by‑step evolution—from SQL tuning and caching to read‑write splitting, high‑availability setups, and finally vertical or horizontal sharding—while warning against premature distribution.

MySQLRead-Write Splittingdistributed systems
0 likes · 15 min read
When to Scale MySQL: From Single Instance to Distributed Architecture
Golang Shines
Golang Shines
Jul 25, 2026 · Operations

Choosing the Right Load Balancer: LVS, Nginx, HAProxy or F5 Explained

When a single server can no longer handle traffic, this guide walks through the four most common load‑balancing solutions—LVS, Nginx, HAProxy and F5—detailing their architectures, configuration steps, scheduling algorithms, pros and cons, and how to pick the best fit for different production scenarios.

F5HAProxyLVS
0 likes · 38 min read
Choosing the Right Load Balancer: LVS, Nginx, HAProxy or F5 Explained
Xiaolin Talks Programming
Xiaolin Talks Programming
Jul 24, 2026 · Backend Development

High-Availability Distributed Task Scheduling with Spring Boot & PowerJob: Production Hardening Guide

This guide shares a year of production experience migrating from basic @Scheduled to PowerJob for high-availability distributed task scheduling, covering Spring Boot integration, DAG workflow orchestration, MapReduce sharding, lease-based failover, JVM tuning, metadata DB protection, and idempotency patterns.

DAG WorkflowDistributed Task SchedulingMapReduce
0 likes · 18 min read
High-Availability Distributed Task Scheduling with Spring Boot & PowerJob: Production Hardening Guide
Ray's Galactic Tech
Ray's Galactic Tech
Jul 21, 2026 · Backend Development

14 Painful Spring Boot Cache Pitfalls and How to Build a High‑Availability Architecture

This article walks through 14 real‑world failure scenarios of Spring Boot distributed caching, explains why high cache hit rates are misleading, and provides concrete analysis, code samples, and step‑by‑step recommendations for designing a resilient cache layer that isolates faults, handles hot keys, and ensures data consistency across Redis, local caches, and databases.

Cache invalidationKubernetesPerformance
0 likes · 37 min read
14 Painful Spring Boot Cache Pitfalls and How to Build a High‑Availability Architecture
Subtle Storm
Subtle Storm
Jul 20, 2026 · Databases

How Database Disaster Recovery Solutions Evolve: From Tape Backups to Multi-Active Geo-Replication

The article explains why disaster‑recovery is essential for architects, defines backup versus DR, introduces RPO/RTO, examines key pitfalls such as sync‑async trade‑offs and split‑brain, compares cold, warm, hot and active‑active setups, and traces the technical evolution from manual tape copies to modern consensus‑based multi‑active geo‑replication.

Multi-ActiveRPORTO
0 likes · 8 min read
How Database Disaster Recovery Solutions Evolve: From Tape Backups to Multi-Active Geo-Replication
IT Learning Made Simple
IT Learning Made Simple
Jul 20, 2026 · Backend Development

Key Takeaways from “Architecture Is the Future”: Scalable Web Architecture Principles

The article distills the core ideas of the book “Architecture Is the Future”, explaining why scalability is essential for modern web services and presenting eight design principles—horizontal scaling, load balancing, fault‑tolerance, data sharding, caching, asynchronous processing, monitoring, and automation—along with organizational patterns, capacity‑planning formulas, performance‑optimization steps, and high‑availability strategies.

CachingPerformance OptimizationWeb Scaling
0 likes · 11 min read
Key Takeaways from “Architecture Is the Future”: Scalable Web Architecture Principles
Golang Shines
Golang Shines
Jul 20, 2026 · Cloud Native

7 Golden Rules for Building High‑Availability Cloud‑Native Go Services (Production‑Proven)

This article presents a step‑by‑step guide to building highly available cloud‑native Go systems, covering graceful error handling, structured logging, minimal dependencies, concurrency control, health checks, Raft‑based replication, timeout/retry strategies, circuit breaking, rate limiting, observability with Zap, Loki, Prometheus, OpenTelemetry, and future architectural directions.

Gocloud-nativedistributed systems
0 likes · 18 min read
7 Golden Rules for Building High‑Availability Cloud‑Native Go Services (Production‑Proven)
CodeSmart Hoops
CodeSmart Hoops
Jul 19, 2026 · Interview Experience

Test Your Grafana Knowledge: 8 Interview Questions with Answers

This article provides a comprehensive Grafana guide covering core concepts, dashboard design principles, panel types, variable templating, alerting strategies, provisioning as code, high‑availability setup, and a multi‑region monitoring screen design, each illustrated with concrete examples and configuration snippets.

GrafanaPrometheusalerting
0 likes · 26 min read
Test Your Grafana Knowledge: 8 Interview Questions with Answers
Java Tech Workshop
Java Tech Workshop
Jul 17, 2026 · Backend Development

Production-Ready WebSocket Connection Pool for Real-Time Market Data with Load Balancing

The article analyzes the fatal issues of using raw WebSocket clients for high‑frequency market feeds and presents a production‑grade, reusable connection‑pool design that adds rate limiting, automatic reconnection, heartbeat, hash‑based load balancing, fault isolation and full lifecycle management to support stable delivery of hundreds of thousands of subscriptions.

Hash ShardingJavaReal-time Data
0 likes · 22 min read
Production-Ready WebSocket Connection Pool for Real-Time Market Data with Load Balancing
Mike Chen Rui
Mike Chen Rui
Jul 16, 2026 · Operations

Inside Alibaba’s City‑Level Active‑Active Architecture: A Complete Guide

The article explains Alibaba’s city‑level active‑active architecture, where two independent data centers in the same city simultaneously serve traffic, automatically fail over on outage, and are organized into traffic, application, and data layers with detailed design patterns.

Active-ActiveAlibabaApplication Layer
0 likes · 4 min read
Inside Alibaba’s City‑Level Active‑Active Architecture: A Complete Guide
IT Learning Made Simple
IT Learning Made Simple
Jul 14, 2026 · Backend Development

Architecture Lessons Learned from Real‑World Failures

The article shares four real‑world failure cases—over‑splitting into microservices, skipping performance testing, accumulating technical debt, and relying on a single‑point database—to illustrate why careful architectural decisions, thorough testing, debt repayment, and high‑availability design are essential for sustainable software systems.

high availabilitymicroservicesperformance testing
0 likes · 9 min read
Architecture Lessons Learned from Real‑World Failures
Raymond Ops
Raymond Ops
Jul 13, 2026 · Operations

Scaling Prometheus to Thousands of Nodes with Thanos: Architecture, Storage, and HA Practices

The article analyzes the storage, query performance, high‑availability, and data‑loss challenges of running Prometheus on a 1,000‑node Kubernetes cluster and demonstrates how a Thanos‑based architecture—Sidecar, Query, Store Gateway, Compactor, Receiver, and object‑storage back‑ends—can be designed, tuned, and operated to achieve horizontal scalability, efficient down‑sampling, and reliable fault recovery.

KubernetesObject StoragePrometheus
0 likes · 35 min read
Scaling Prometheus to Thousands of Nodes with Thanos: Architecture, Storage, and HA Practices
IoT Full-Stack Technology
IoT Full-Stack Technology
Jul 13, 2026 · Databases

Is Sharding Dead? A Technical Comparison with NewSQL Databases

The article objectively compares middleware‑based sharding with NewSQL distributed databases, examining distributed transactions, CAP constraints, HA, scaling, storage engines, maturity, and ecosystem, and offers a decision framework for choosing the appropriate architecture based on concrete requirements.

CAP theoremNewSQLdatabase architecture
0 likes · 18 min read
Is Sharding Dead? A Technical Comparison with NewSQL Databases
Linyb Geek Road
Linyb Geek Road
Jul 12, 2026 · Operations

Designing a High‑Availability Architecture: Core Principles and Practices

This article outlines the essential principles for building a high‑availability system, covering cluster and distributed designs, fault‑tolerance, reliable hardware, disaster recovery, monitoring, security, capacity planning, and automated scaling to achieve optimal performance and resilience.

automationcapacity-planningcluster architecture
0 likes · 6 min read
Designing a High‑Availability Architecture: Core Principles and Practices
IT Learning Made Simple
IT Learning Made Simple
Jul 10, 2026 · Backend Development

What a Forced Resignation Reveals About Single‑Point Failure and High‑Availability

The viral workplace drama where a key engineer is forced out serves as a vivid case study, showing how relying on a single technical pillar creates a single‑point failure that can bring an entire online business down, and illustrating the essential practices of layered architecture and high‑availability design.

IT learninghigh availabilitysingle point failure
0 likes · 7 min read
What a Forced Resignation Reveals About Single‑Point Failure and High‑Availability
Ctrip Technology
Ctrip Technology
Jul 10, 2026 · Cloud Native

How Ctrip Scaled Karmada to Handle 200 GB Memory Peaks and Hundreds of Thousands of Pods

This article details Ctrip's migration from Kubefed to Karmada for multi‑cluster governance, describing the architectural evolution, production deployment for high‑availability and smooth cross‑cluster pod migration, and the performance and memory optimizations required to support over 200 GB of control‑plane memory usage and tens of thousands of Pods.

Control Plane OptimizationFederationKarmada
0 likes · 18 min read
How Ctrip Scaled Karmada to Handle 200 GB Memory Peaks and Hundreds of Thousands of Pods
CTO Full-Stack Academy
CTO Full-Stack Academy
Jul 10, 2026 · Backend Development

Designing a Scalable Email Management System: Core Principles and Implementation

The article presents a comprehensive, step‑by‑step design for an enterprise‑grade email platform that handles billions of daily messages, detailing goals, architectural patterns, microservice decomposition, storage tiers, queue‑based throttling, spam protection, and high‑availability strategies.

email systemhigh availabilitymicroservices
0 likes · 18 min read
Designing a Scalable Email Management System: Core Principles and Implementation
YiSu Grain
YiSu Grain
Jul 9, 2026 · Backend Development

Mapping Database & Architecture Patterns onto an E‑Commerce High‑Concurrency Diagram

This article reviews weeks 8‑13 of a system‑architecture course—covering indexes, ACID, MVCC, high availability, performance tuning, and case‑study templates—and shows how to combine those concepts into a complete e‑commerce high‑concurrency solution with caching, load‑balancing, async processing, database optimization, HA clustering, and concurrency control.

Cachingdatabase optimizatione-commerce
0 likes · 18 min read
Mapping Database & Architecture Patterns onto an E‑Commerce High‑Concurrency Diagram
YiSu Grain
YiSu Grain
Jul 9, 2026 · Fundamentals

How to Write Score‑Winning Answers for Architecture Case Questions

The article explains why simply listing technical terms in a software‑exam case study earns no points and provides a step‑by‑step method—identifying problems, mapping them to architectural patterns, and phrasing solutions as concrete, business‑focused sentences that score well.

CachingPerformance Optimizationarchitecture-design
0 likes · 14 min read
How to Write Score‑Winning Answers for Architecture Case Questions
Architecture & Thinking
Architecture & Thinking
Jul 9, 2026 · Backend Development

Prevent Cache Avalanche with Multi‑Level Caffeine + Redis: High‑Availability Design

The article explains how combining a local Caffeine cache with a Redis cluster in a three‑tier architecture can protect high‑traffic distributed systems from cache avalanche, detailing expiration strategies, hot‑cold data separation, fault‑tolerant fallback, consistency handling, performance benchmarks, and practical pitfalls.

CaffeineJavaMulti-level Cache
0 likes · 18 min read
Prevent Cache Avalanche with Multi‑Level Caffeine + Redis: High‑Availability Design
Yumin Fish Harvest
Yumin Fish Harvest
Jul 7, 2026 · Databases

Redis High‑Availability Deep Dive: Master‑Slave Replication, Sentinel, and Split‑Brain Protection

This article explains why a single‑node Redis deployment is a single‑point‑of‑failure and walks through building a highly available Redis cluster using master‑slave replication, Sentinel monitoring and automatic failover, split‑brain prevention, production deployment guidelines, common pitfalls, and client‑side connection strategies.

ConfigurationRedisReplication
0 likes · 36 min read
Redis High‑Availability Deep Dive: Master‑Slave Replication, Sentinel, and Split‑Brain Protection
dbaplus Community
dbaplus Community
Jul 6, 2026 · Cloud Native

Why Companies Are Switching from Kubernetes to K3s: Simplicity, Stability, and Low Overhead

The article explains how K3s, a lightweight, production‑grade Kubernetes distribution, reduces component count, memory usage, and installation complexity, making it ideal for small‑to‑medium enterprises, edge, IoT, and AI projects, while still offering API compatibility and optional high‑availability, and compares its trade‑offs with full‑size Kubernetes.

InstallationKubernetesedge computing
0 likes · 7 min read
Why Companies Are Switching from Kubernetes to K3s: Simplicity, Stability, and Low Overhead
Cloud Architecture
Cloud Architecture
Jul 2, 2026 · Databases

MySQL Containerization vs Host Installation: Production‑Grade Selection Framework

The article explains that the real challenge is not merely running MySQL but placing it in the right resource model, and it provides a four‑dimensional decision framework—performance ceiling, stability floor, automation level, and organizational maturity—to guide when to use host‑installed MySQL, single‑node containers, or full Kubernetes deployment, illustrated with concrete resource analyses, architecture diagrams, configuration examples, pitfalls, checklists, and an evolution roadmap.

ContainerizationKubernetesMySQL
0 likes · 35 min read
MySQL Containerization vs Host Installation: Production‑Grade Selection Framework
Subtle Storm
Subtle Storm
Jul 1, 2026 · Backend Development

How to Tackle the “Three Highs” of Internet Systems Without Burning Out

The article analyzes the intertwined challenges of high concurrency, high performance, and high availability in internet services, explains why they cannot all be maximized simultaneously, and presents concrete architectural tactics—partitioning, caching, async processing, redundancy, and CAP trade‑offs—to achieve a balanced, resilient system.

CAP theoremCachingHigh Performance
0 likes · 7 min read
How to Tackle the “Three Highs” of Internet Systems Without Burning Out
Code Farming
Code Farming
Jul 1, 2026 · Databases

Redis Core Principles Explained with Four Diagrams

This article breaks down Redis’s core mechanisms—including its single‑threaded performance tricks, AOF and RDB persistence designs, and the evolution of high‑availability from replication to Sentinel and Cluster—using four clear diagrams to help readers master the system.

AOFPersistenceRDB
0 likes · 6 min read
Redis Core Principles Explained with Four Diagrams
Raymond Ops
Raymond Ops
Jun 27, 2026 · Operations

Hands‑On DNS Ops: Deploy BIND and CoreDNS with Full Troubleshooting Guide

This comprehensive guide walks you through DNS fundamentals, compares BIND, CoreDNS, PowerDNS and Unbound, provides step‑by‑step deployment scripts for BIND 9.20 and CoreDNS 1.12, explains DNSSEC configuration, caching optimizations, security hardening, high‑availability designs, monitoring, backup and recovery procedures, and advanced troubleshooting techniques.

BINDCoreDNSDNS
0 likes · 43 min read
Hands‑On DNS Ops: Deploy BIND and CoreDNS with Full Troubleshooting Guide
Xiaolin Talks Programming
Xiaolin Talks Programming
Jun 27, 2026 · Backend Development

Spring Boot WebSocket Clustering: State Decoupling & Routing at Scale

This article details production-hardened patterns for scaling Spring Boot WebSocket clusters, covering protocol trade-offs, STOMP heartbeat alignment, Redis-based routing, offline message compensation, reconnection backoff, idempotency layers, Undertow tuning, memory leak diagnostics, Nginx proxy pitfalls, split-brain defense, and graceful shutdown — all grounded in real incidents.

ClusteringNginxRedis Pub/Sub
0 likes · 21 min read
Spring Boot WebSocket Clustering: State Decoupling & Routing at Scale
Long Ge's Treasure Box
Long Ge's Treasure Box
Jun 26, 2026 · Operations

Designing High‑Availability Systems: Multi‑Active Architectures, Failover, Monitoring, and SLO/SLI

This article explains how to build highly available services by comparing single‑datacenter, same‑city active‑active, two‑city three‑center, and global multi‑active architectures, then details health‑check mechanisms, automatic failover workflows, Prometheus‑Grafana monitoring, and SLO/SLI error‑budget management with concrete code examples.

SLISLOfailover
0 likes · 16 min read
Designing High‑Availability Systems: Multi‑Active Architectures, Failover, Monitoring, and SLO/SLI
Random Bulletin
Random Bulletin
Jun 23, 2026 · Backend Development

Choosing the Right Replication Factor for 10M QPS Message Queues: From Dual to Multi‑Replica

The article walks through why a single replica is insufficient for million‑scale message queues, explains the latency and data‑loss trade‑offs of dual‑replica sync and async modes, shows how three‑replica majority voting becomes the sweet spot, and then details the engineering considerations for scaling to five or more replicas across availability zones and regions.

AZ awarenessKafkaPulsar
0 likes · 18 min read
Choosing the Right Replication Factor for 10M QPS Message Queues: From Dual to Multi‑Replica
Cloud Architecture
Cloud Architecture
Jun 22, 2026 · Databases

Opening the Hood of PostgreSQL WAL: From Transaction Logs to Production‑Grade HA, Replication, and Performance Tuning

This article dissects PostgreSQL's Write‑Ahead Logging (WAL), explaining how it underpins transaction durability, replication, CDC, point‑in‑time recovery, and commit latency, and provides a step‑by‑step guide to architecture design, parameter tuning, troubleshooting, containerization, backup strategies, and operational checklists for production‑grade deployments.

CDCPostgreSQLReplication
0 likes · 37 min read
Opening the Hood of PostgreSQL WAL: From Transaction Logs to Production‑Grade HA, Replication, and Performance Tuning
Raymond Ops
Raymond Ops
Jun 20, 2026 · Operations

Eliminate Monitoring Blind Spots: Hands‑On Enterprise‑Grade Prometheus + Grafana Deployment

This comprehensive guide walks you through the end‑to‑end setup of a production‑grade Prometheus and Grafana monitoring stack, covering architecture choices, installation steps, configuration details, high‑availability designs, performance tuning, security hardening, troubleshooting, backup strategies, and best‑practice recommendations.

GrafanaKubernetesPrometheus
0 likes · 49 min read
Eliminate Monitoring Blind Spots: Hands‑On Enterprise‑Grade Prometheus + Grafana Deployment
Java Architect Handbook
Java Architect Handbook
Jun 18, 2026 · Cloud Native

Designing an Enterprise-Grade Message Push Architecture: A Deep Dive

The article outlines the evolution from isolated push modules to a unified framework and finally a dedicated push service, detailing functional and non‑functional requirements, component responsibilities, priority handling, and a scalable micro‑service architecture for enterprise notifications.

Message PushNotification Serviceenterprise architecture
0 likes · 15 min read
Designing an Enterprise-Grade Message Push Architecture: A Deep Dive
Raymond Ops
Raymond Ops
Jun 17, 2026 · Databases

Redis Sentinel Mode Explained: Automatic Failure Detection and Master‑Slave Switching in Practice

This guide walks through Redis Sentinel’s architecture, explains subjective and objective down states, details the leader election and failover workflow, shows step‑by‑step configuration of a three‑node Sentinel cluster, client integration in Python and Java, and provides best‑practice recommendations, monitoring metrics, and troubleshooting tips.

ConfigurationJavaPython
0 likes · 27 min read
Redis Sentinel Mode Explained: Automatic Failure Detection and Master‑Slave Switching in Practice
Raymond Ops
Raymond Ops
Jun 17, 2026 · Operations

Enterprise Monitoring with Prometheus: Rule Hierarchy and Alertmanager Notification Orchestration

This guide explains how to turn a fully built Prometheus monitoring system into a closed‑loop alerting solution by designing layered PromQL rules, configuring Alertmanager routing, grouping, inhibition and silencing, integrating DingTalk and WeChat webhooks, and applying best‑practice performance, security, high‑availability, and troubleshooting techniques.

AlertmanagerDevOpsKubernetes
0 likes · 34 min read
Enterprise Monitoring with Prometheus: Rule Hierarchy and Alertmanager Notification Orchestration
Mike Chen's Internet Architecture
Mike Chen's Internet Architecture
Jun 16, 2026 · Operations

How Alibaba’s Two‑Region Three‑Center Design Achieves 99.99% Availability

The article explains Alibaba’s “two‑region three‑center” architecture, detailing how geographically separated primary, backup, and disaster‑recovery data centers work together to provide financial‑grade high availability and protect against single‑site failures or regional catastrophes.

AlibabaData Center Architecturedisaster recovery
0 likes · 3 min read
How Alibaba’s Two‑Region Three‑Center Design Achieves 99.99% Availability
Coder Trainee
Coder Trainee
Jun 14, 2026 · Artificial Intelligence

Production‑Ready AI Agent Architecture: High Availability, Asynchrony, Caching, Cost & Security

After mastering core AI Agent capabilities, this article shows how to transform a prototype into a production‑grade service by covering a full architecture overview, stateless design, health‑check and graceful shutdown, asynchronous task queues, multi‑level caching, token‑cost optimization, model fallback, input/output filtering, rate limiting, monitoring, and deployment recommendations for different scales.

AI agentCachingProduction Architecture
0 likes · 15 min read
Production‑Ready AI Agent Architecture: High Availability, Asynchrony, Caching, Cost & Security
Mike Chen's Internet Architecture
Mike Chen's Internet Architecture
Jun 12, 2026 · Industry Insights

Inside Alibaba’s Same‑City Active‑Active Architecture: A Complete Visual Guide

The article breaks down Alibaba’s same‑city active‑active high‑availability architecture, detailing its four design layers—traffic scheduling, stateless application services, data replication, and operational automation—while illustrating how each component ensures continuous service during data‑center failures.

Active-ActiveAlibabaStateless Architecture
0 likes · 5 min read
Inside Alibaba’s Same‑City Active‑Active Architecture: A Complete Visual Guide
Architect's Guide
Architect's Guide
Jun 12, 2026 · Operations

Common Disaster Recovery Models and How to Choose Them

The article outlines the main disaster‑recovery architectures—city‑level, remote, two‑site three‑center, and active‑active data centers—explains their characteristics, compares costs and performance, and presents key selection metrics such as RPO, RTO, disaster radius and ROI, illustrated with Huawei and ZTE case studies.

RPORTOarchitecture
0 likes · 13 min read
Common Disaster Recovery Models and How to Choose Them
Cloud Architecture
Cloud Architecture
Jun 11, 2026 · Backend Development

RocketMQ Storage HA Deep Dive: CommitLog Mechanics to Production Controller Failover

This article analyzes RocketMQ’s storage high‑availability, detailing the CommitLog, ConsumeQueue, flushing and replication mechanisms, comparing traditional master‑slave, DLedger and Controller modes, and provides engineering configurations, code examples, capacity planning, Kubernetes deployment, monitoring and recovery practices for production‑grade fault tolerance.

CommitLogControllerDLedger
0 likes · 42 min read
RocketMQ Storage HA Deep Dive: CommitLog Mechanics to Production Controller Failover
Xiao Liu Lab
Xiao Liu Lab
Jun 11, 2026 · Operations

Ops Engineer Core Skills: From Basic Commands to High‑Availability Architecture

This article provides a comprehensive roadmap for operations engineers, covering essential Linux commands, core system concepts, service principles, fault‑diagnosis methods, high‑availability architecture designs, data security, backup strategies, performance tuning, and automation scripts to handle both single‑machine and large‑scale cluster environments.

KubernetesLinuxMySQL
0 likes · 13 min read
Ops Engineer Core Skills: From Basic Commands to High‑Availability Architecture
Java Tech Workshop
Java Tech Workshop
Jun 8, 2026 · Databases

Advanced SpringBoot Read‑Write Splitting: Master‑Slave Switching and Automatic Failover

In high‑concurrency internet architectures, a MySQL master‑slave setup with read‑write splitting is the baseline for high availability, but static routing suffers from node failures and lag; this article explains how ShardingSphere provides health checks, auto‑failover, load‑balancing, and degradation to achieve resilient read‑write separation.

MySQLRead-Write SplittingShardingSphere
0 likes · 14 min read
Advanced SpringBoot Read‑Write Splitting: Master‑Slave Switching and Automatic Failover
Coder Trainee
Coder Trainee
Jun 6, 2026 · Backend Development

Spring Cloud Message‑Driven Part 5: High‑Availability RocketMQ Deployment & Message Tracing

This tutorial walks through deploying a highly available RocketMQ cluster with Docker Compose, configuring master‑slave brokers, enabling message tracing, integrating Prometheus‑Grafana monitoring, setting up Spring Boot HA properties, applying performance tweaks, validating failover, and troubleshooting common issues.

Docker ComposeGrafanaMessage Tracing
0 likes · 16 min read
Spring Cloud Message‑Driven Part 5: High‑Availability RocketMQ Deployment & Message Tracing
Raymond Ops
Raymond Ops
Jun 5, 2026 · Operations

Dual‑Master Nginx + Keepalived Architecture: Eliminate Single Points of Failure

This guide walks through building a dual‑master Nginx + Keepalived high‑availability setup that doubles resource utilization, removes the idle‑backup drawback of traditional active‑passive designs, and provides step‑by‑step configuration, health‑check scripts, failover testing, best‑practice tips, and troubleshooting procedures.

KeepalivedLinuxNginx
0 likes · 33 min read
Dual‑Master Nginx + Keepalived Architecture: Eliminate Single Points of Failure
ITPUB
ITPUB
Jun 4, 2026 · Backend Development

How to Ensure High Availability When Third‑Party Services Fail?

The article explains how to protect a system from unstable third‑party APIs by building an isolated defense layer that offers a unified abstraction, client‑side rate limiting and retry, comprehensive observability, and mock testing, and shows how to present these solutions in technical interviews.

Circuit Breakingclient-side rate limitinghigh availability
0 likes · 21 min read
How to Ensure High Availability When Third‑Party Services Fail?
DeepNoMind
DeepNoMind
May 31, 2026 · Databases

Inside PayPal’s JunoDB: A High‑Performance Distributed KV Store

The article examines PayPal’s open‑source JunoDB, a high‑performance distributed key‑value store built in Go, detailing its motivation, three‑tier proxy architecture, sharding strategy, quorum‑based consistency, aggressive replication, security mechanisms, real‑world PayPal use cases, and the path to open‑source release.

GoJunoDBPayPal
0 likes · 11 min read
Inside PayPal’s JunoDB: A High‑Performance Distributed KV Store
MaGe Linux Operations
MaGe Linux Operations
May 20, 2026 · Operations

How to Choose Among the Four Common Load‑Balancing Solutions: LVS, Nginx, HAProxy or F5

This article explains why single‑server capacity is limited, lists typical load‑balancing problems, and provides a detailed comparison of four mainstream solutions—LVS, Nginx, HAProxy, and F5—covering their principles, architectures, configuration steps, pros, cons, suitable scenarios, a decision‑tree guide, common fault‑diagnosis procedures, and production‑risk warnings.

F5HAProxyLVS
0 likes · 38 min read
How to Choose Among the Four Common Load‑Balancing Solutions: LVS, Nginx, HAProxy or F5
Cloud Architecture
Cloud Architecture
May 19, 2026 · Operations

RabbitMQ High‑Availability Cluster: Theory, Architecture, and Production Troubleshooting

This article explains why RabbitMQ failures can cascade through a micro‑service system, details the underlying HA mechanisms such as quorum queues, presents a layered production architecture with concrete Spring Boot code, outlines a step‑by‑step troubleshooting workflow, and shares best‑practice checklists for scaling, Kubernetes deployment, and migration from classic mirrored queues.

KubernetesQuorum QueueRabbitMQ
0 likes · 52 min read
RabbitMQ High‑Availability Cluster: Theory, Architecture, and Production Troubleshooting
Architects' Tech Alliance
Architects' Tech Alliance
May 16, 2026 · Industry Insights

Designing a 2026 Ultra‑Large Green AI Data Center: Full Infrastructure Blueprint

This article presents a comprehensive 2026 design plan for an ultra‑large green AI data center with 5,000 cabinets, 150 MW IT load, and 200 MW capacity, detailing market drivers, core metrics, six design principles, site and power architecture, liquid‑cooling, networking, security, and AI‑driven autonomous operations.

2026 designAI data centerDCIM
0 likes · 5 min read
Designing a 2026 Ultra‑Large Green AI Data Center: Full Infrastructure Blueprint
AI Agent Super App
AI Agent Super App
May 13, 2026 · Operations

Server Virtualization Deep Dive: Feature Comparison of VMware, KVM, Proxmox and Practical High‑Availability

This comprehensive guide walks through server virtualization fundamentals, compares major hypervisors such as VMware vSphere, KVM, Xen, Proxmox VE and Hyper‑V, and then details Linux‑level monitoring, performance tuning, backup strategies, and cross‑node high‑availability solutions for production environments.

KVMProxmoxVMware
0 likes · 24 min read
Server Virtualization Deep Dive: Feature Comparison of VMware, KVM, Proxmox and Practical High‑Availability
Subtle Storm
Subtle Storm
May 12, 2026 · Backend Development

A Ready‑to‑Use Template for Scoring High on the System Architecture Designer Exam

This article provides a comprehensive, step‑by‑step template for a high‑scoring system architecture design paper, detailing project background, challenges, six‑stage ABSD design, microservice migration with Spring Cloud Alibaba and Kubernetes, performance metrics, high‑availability safeguards, and lessons learned.

Design PatternsKubernetesPerformance Optimization
0 likes · 10 min read
A Ready‑to‑Use Template for Scoring High on the System Architecture Designer Exam
Ops Community
Ops Community
May 9, 2026 · Operations

Achieve Seamless Nginx High Availability with Keepalived: A Practical Guide

This article walks through building a simple, cost‑effective high‑availability solution for Nginx using Keepalived’s VRRP‑based VIP failover, covering environment setup, configuration of master and backup nodes, health‑check scripts, testing procedures, troubleshooting tips, and rollback steps.

KeepalivedLinuxNginx
0 likes · 29 min read
Achieve Seamless Nginx High Availability with Keepalived: A Practical Guide
JD Tech
JD Tech
May 8, 2026 · Databases

Engineering Wisdom Behind High‑Availability Architecture for E‑Commerce Storage Layers

The article analyzes how to design a high‑availability architecture for large‑scale e‑commerce systems, detailing layered risk isolation, stateful storage strategies for flow and state data, unified document‑ID routing, multi‑replica databases, multi‑datacenter synchronization, and real‑world JD case studies that demonstrate elastic scaling and disaster recovery.

Database ReplicationDistributed Architecturee-commerce
0 likes · 17 min read
Engineering Wisdom Behind High‑Availability Architecture for E‑Commerce Storage Layers
Linyb Geek Road
Linyb Geek Road
May 7, 2026 · Backend Development

How to Ensure High Availability When Third‑Party Services Keep Failing – An Interview‑Ready Guide

The article explains how to design a defensive layer that abstracts third‑party calls, implements client‑side rate limiting, retries, circuit breaking, observability, and mock testing, and shows how to present these practices effectively during a system‑design interview.

Circuit BreakerInterview Preparationhigh availability
0 likes · 21 min read
How to Ensure High Availability When Third‑Party Services Keep Failing – An Interview‑Ready Guide
Linyb Geek Road
Linyb Geek Road
May 7, 2026 · Operations

A Decade of E‑Commerce Ops: How to Prevent System Outages and Ensure High Availability

The article outlines why e‑commerce systems fail, presents a four‑layer high‑availability defense—including load balancing, service isolation, data protection, and fallback mechanisms—plus concrete monitoring, alerting, and emergency response practices illustrated with real‑world scenarios and code samples.

database backupdisaster recoverye-commerce
0 likes · 6 min read
A Decade of E‑Commerce Ops: How to Prevent System Outages and Ensure High Availability
dbaplus Community
dbaplus Community
Apr 28, 2026 · Backend Development

Designing High‑Availability for Unreliable Third‑Party Services

When downstream APIs are unstable and slow, this article walks through building a dedicated defensive layer that provides a unified abstraction, client‑side governance (rate limiting, retries with idempotency checks), comprehensive observability, and mock‑based testing to keep your system highly available and interview‑ready.

Circuit BreakerThird-party Integrationhigh availability
0 likes · 22 min read
Designing High‑Availability for Unreliable Third‑Party Services
Java Backend Full-Stack
Java Backend Full-Stack
Apr 27, 2026 · Databases

Proven Redis Tuning Techniques for Production Environments

This article compiles practical, interview‑ready Redis tuning tips—from strict memory limits and eviction policies to avoiding big keys, hot keys, slow commands, and optimizing persistence, networking, and high‑availability settings—so you can confidently handle Redis performance questions in real‑world deployments.

ConfigurationRedishigh availability
0 likes · 9 min read
Proven Redis Tuning Techniques for Production Environments
Software Engineering 3.0 Era
Software Engineering 3.0 Era
Apr 22, 2026 · Operations

Is This the Longest Service Outage Ever Recorded?

The article examines the航旅纵横 app’s service disruption that began at 12:30 PM on April 21 and lasted until 7:47 AM on April 22—over 19 hours—questioning whether this duration ranks among the longest outages ever, citing official posts, AI‑generated rankings, and a reminder that high‑availability depends on disciplined engineering rather than tools.

AIdowntimehigh availability
0 likes · 3 min read
Is This the Longest Service Outage Ever Recorded?
Cloud Architecture
Cloud Architecture
Apr 17, 2026 · Databases

From Crashing Databases to Zero Data Loss: MySQL Replication Evolution

The article examines common MySQL replication failures, explains why traditional position‑based replication is insufficient, and walks through a step‑by‑step evolution from file/position to GTID, asynchronous to semi‑synchronous, single‑threaded to parallel apply, culminating in a production‑grade architecture that achieves near‑zero data loss.

GTIDMySQLParallel Replication
0 likes · 35 min read
From Crashing Databases to Zero Data Loss: MySQL Replication Evolution
Lobster Programming
Lobster Programming
Apr 15, 2026 · Databases

Choosing the Right Redis Architecture: From Single Node to Cluster

This article reviews the main Redis deployment options—including single‑node, master‑slave with Sentinel, sharding via consistent hashing, and Redis Cluster—explaining their advantages, high‑availability mechanisms, scalability limits, and recommending suitable scenarios for each architecture.

Redisclusterdeployment
0 likes · 7 min read
Choosing the Right Redis Architecture: From Single Node to Cluster