Tagged articles

distributed systems

2274 articles · Page 1 of 23
Data Bricklaying Diary
Data Bricklaying Diary
Oct 7, 2026 · Backend Development

Agent Error Recovery: Why Retries Alone Fail — Idempotency & Reconciliation

This article explains why automatic retries after agent timeouts can duplicate business actions, and details how to distinguish call failures from executed side effects, implement execution-side idempotency guarantees, reconcile using original action identifiers, and define safe recovery rules with human escalation when outcomes remain unknown.

Agent systemsIdempotencydistributed systems
0 likes · 17 min read
Agent Error Recovery: Why Retries Alone Fail — Idempotency & Reconciliation
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Oct 3, 2026 · Artificial Intelligence

DeepSeek DSec: Running 380K Sandboxes with 50x Overcommit for Agent Training

DeepSeek's DSec infrastructure supports millions of agent sandboxes through layered environments, on-demand image loading via 3FS, memory sharing with virtio-pmem/DAX, CPU scheduling, and trajectory forking, achieving 50x resource overcommit while addressing security challenges like agent-discovered vulnerabilities.

3FSAgent trainingAppArmor
0 likes · 12 min read
DeepSeek DSec: Running 380K Sandboxes with 50x Overcommit for Agent Training
Random Bulletin
Random Bulletin
Oct 1, 2026 · Backend Development

Fault Domain Design: Turning Blast Radius from 100% into a Tunable 1/N Parameter

The article presents a layered fault-domain strategy—physical anti-affinity, logical isolation (sharding, cluster groups, swimlanes, bulkheads), cell-based architecture, chaos-engineering validation, and quantitative governance metrics—to shrink the blast radius of a ten-million-QPS system from a fixed 100% to a controllable 1/N design parameter.

KubernetesSLOanti-affinity
0 likes · 25 min read
Fault Domain Design: Turning Blast Radius from 100% into a Tunable 1/N Parameter
Random Bulletin
Random Bulletin
Sep 30, 2026 · Backend Development

Cross-Region Disaster Recovery: From Backup to Multi-Active at 10M QPS

This article systematically dissects the engineering evolution from disaster backup to multi-active architecture, covering fault domain isolation, RTO/RPO/recovery capacity metrics, data replication trade-offs, traffic switching state machines, business unitization, conflict convergence, and a phased adoption roadmap for systems operating at ten million QPS.

Conflict ResolutionRTO RPO recovery capacitybusiness unitization
0 likes · 39 min read
Cross-Region Disaster Recovery: From Backup to Multi-Active at 10M QPS
dbaplus Community
dbaplus Community
Sep 29, 2026 · Industry Insights

From Monoliths to AI Agents: Architecture Evolution's Two Patterns and New Challenges

This article traces software architecture evolution from monoliths through primitive distributed systems, SOA, microservices, and cloud-native, highlighting two recurring patterns — finer decoupling and stronger fault isolation — and a shift from zero-failure goals to designing for resilience, while outlining four unprecedented challenges AI agents introduce: semantic hallucinations, stateful context, dynamic orchestration, and observability gaps.

AI agentscloud nativedistributed systems
0 likes · 23 min read
From Monoliths to AI Agents: Architecture Evolution's Two Patterns and New Challenges
Ops Development & AI Practice
Ops Development & AI Practice
Sep 25, 2026 · Interview Experience

Forgetting Details ≠ Incompetence: Senior Engineers' Cognitive Compression Playbook

This article explains why senior engineers naturally forget micro-details through the brain's lossy compression, and provides a three-part framework — pre-interview 'deep-water nails' (real failure cases, core parameters, personal artifacts) and three on-the-spot defense frameworks (deduction, methodology anchoring, pivoting) — to demonstrate unforgeable engineering depth in high-stakes technical interviews.

Goarchitecture reviewcognitive compression
0 likes · 18 min read
Forgetting Details ≠ Incompetence: Senior Engineers' Cognitive Compression Playbook
Data Bricklaying Diary
Data Bricklaying Diary
Sep 25, 2026 · Backend Development

Beyond Retry: Designing Fault Tolerance from Failure Models to Disaster Recovery

This article explains why retry alone is insufficient for fault tolerance, detailing how to classify failures via fault models, apply appropriate mechanisms like timeouts and circuit breakers, design recovery paths targeting state convergence, define RTO/RPO for disaster recovery, and validate all assumptions through fault injection drills.

IdempotencyRPORTO
0 likes · 23 min read
Beyond Retry: Designing Fault Tolerance from Failure Models to Disaster Recovery
liandk
liandk
Sep 25, 2026 · Backend Development

MQ Production Failure Troubleshooting: Message Loss, Duplicates, Backlog & Dead Letters

This comprehensive guide covers the five critical MQ production failures — message loss, duplicate consumption, massive backlog, consumer hangs, and dead letter queue blocking — with root cause analysis, emergency mitigation steps, and long-term architectural fixes for RocketMQ, Kafka, and RabbitMQ.

KafkaMQMessage Queue
0 likes · 14 min read
MQ Production Failure Troubleshooting: Message Loss, Duplicates, Backlog & Dead Letters
Cloud Architecture
Cloud Architecture
Sep 24, 2026 · Backend Development

Scaling Spring Boot Sign-In to 100M Users: Distributed Architecture Evolution & Production Hardening

This article details the evolution of a Spring Boot daily sign-in system from 10K to 100M users, covering Redis bitmap sharding, Lua atomic operations, command outbox pattern, Kafka async reward processing, idempotency guarantees, fault tolerance strategies, and production-grade observability with real incident postmortems.

BitmapCommand OutboxIdempotency
0 likes · 36 min read
Scaling Spring Boot Sign-In to 100M Users: Distributed Architecture Evolution & Production Hardening
Xiaolin Talks Programming
Xiaolin Talks Programming
Sep 23, 2026 · Backend Development

Spring Boot QR Code Login: State Machines, SSE Push, and Replay Protection for Production

This article details a production-ready QR code login implementation using Spring Boot, covering state machine design with Redis and Lua for atomic transitions, SSE for low-latency status push with polling fallback, HMAC-SHA256 with nonce for replay protection, and Redis Pub/Sub for cross-instance consistency in clustered deployments.

Lua ScriptsQR Code LoginRedis
0 likes · 28 min read
Spring Boot QR Code Login: State Machines, SSE Push, and Replay Protection for Production
LuTiao Programming
LuTiao Programming
Sep 21, 2026 · Backend Development

Distributed Rate Limiting with Redis + Lua: Surviving API Floods in Spring Boot

After an external system hammered a Spring Boot search endpoint causing database connection exhaustion, the author builds a distributed fixed-window rate limiter using Redis and Lua for atomicity, wraps it with an annotation-driven AOP aspect supporting user, IP, API key, and global dimensions, returns proper HTTP 429 with Retry-After, discusses fixed-window limitations versus token bucket and sliding window, and covers fail-open/fail-closed strategies for Redis outages plus monitoring metrics.

AOPFail-OpenFixed Window
0 likes · 17 min read
Distributed Rate Limiting with Redis + Lua: Surviving API Floods in Spring Boot
Random Bulletin
Random Bulletin
Sep 20, 2026 · Backend Development

Automating Fault Localization at 10M QPS: Evidence Chains Over Manual Hunts

This article details how to build automated fault localization for 10M QPS systems by unifying entity identities, aligning timestamps, integrating change records, and applying a four-layer engine—anomaly normalization, temporal correlation, topological pruning, and causal scoring—to converge millions of anomalies into verifiable hypotheses while avoiding correlation-causation pitfalls through counterfactual evidence and phased rollout.

Causal Inferenceautomated troubleshootingdistributed systems
0 likes · 35 min read
Automating Fault Localization at 10M QPS: Evidence Chains Over Manual Hunts
Java Tech Workshop
Java Tech Workshop
Sep 20, 2026 · Backend Development

Order Timeout Auto-Cancellation: RabbitMQ Delayed Queue + Scheduled Task Dual Insurance Pattern

This article details a production-ready dual-insurance pattern for e-commerce order timeout cancellation, combining RabbitMQ delayed message queues for real-time processing with scheduled database scans as a fallback, including Spring Boot implementation code, idempotent cancellation logic, and distributed deployment considerations.

IdempotencyMessage QueueRabbitMQ
0 likes · 16 min read
Order Timeout Auto-Cancellation: RabbitMQ Delayed Queue + Scheduled Task Dual Insurance Pattern
LuTiao Programming
LuTiao Programming
Sep 20, 2026 · Backend Development

How to Guarantee Single Order Processing When Payment Callbacks Fire 10 Times

The article explains how to handle duplicate payment callbacks in Spring Boot using database unique constraints, conditional state updates, an outbox pattern for reliable event publishing, and downstream idempotency, ensuring orders are processed exactly once even if the payment platform sends repeated notifications.

Database ConstraintsIdempotencyOutbox Pattern
0 likes · 17 min read
How to Guarantee Single Order Processing When Payment Callbacks Fire 10 Times
Random Bulletin
Random Bulletin
Sep 19, 2026 · Operations

Second-Level Fault Detection at 10M QPS: Layered Signals & Safe Automation

This article explains how to reduce fault detection latency from minutes to seconds in 10M QPS systems by implementing layered signals, combined evidence detection, distributed judgment, event normalization, and safe automation guardrails, rather than simply increasing sampling frequency.

AlertingSLOdistributed systems
0 likes · 36 min read
Second-Level Fault Detection at 10M QPS: Layered Signals & Safe Automation
Architect's Guide
Architect's Guide
Sep 16, 2026 · Backend Development

Implementing Cross-System Single Sign-On with CAS: A Practical Guide

This article explains how to implement Single Sign-On (SSO) using CAS to unify authentication across multiple company systems, covering session mechanics, cluster session sharing with Redis, and providing complete Spring Boot code examples for both the CAS server and client applications.

AuthenticationCASRedis
0 likes · 16 min read
Implementing Cross-System Single Sign-On with CAS: A Practical Guide
AI Engineering
AI Engineering
Sep 15, 2026 · Operations

Anthropic's CI Crisis: Scaling Test Impact Analysis After AI Wrote 80% of Code

Anthropic's AI-generated code increased CI jobs 25x, overwhelming their Test Impact Analysis service; they iterated through three patches before redesigning with a distributed journal-based architecture that stabilized queue backlog, teaching them to plan for exponential growth and separate state from process.

AI-generated codeAnthropicCI/CD
0 likes · 5 min read
Anthropic's CI Crisis: Scaling Test Impact Analysis After AI Wrote 80% of Code
Architect
Architect
Sep 13, 2026 · Artificial Intelligence

Multi-Agent Consistency: Distributed Systems Challenges Return with Autonomous Agents

The article explores four critical questions for multi-agent consistency: task decomposition rationale, structured handoffs with versioned snapshots, conflict resolution via evidence-based contracts, and verifiable completion criteria. It argues multi-agent systems reintroduce classic distributed systems challenges—identity, leases, idempotency, compensation—and require runtime proofs over model assertions.

Agent ArchitectureConsistencyIdempotency
0 likes · 21 min read
Multi-Agent Consistency: Distributed Systems Challenges Return with Autonomous Agents
Random Bulletin
Random Bulletin
Sep 13, 2026 · Backend Development

Safe Configuration Rollout at Scale: Versioning, Canary, and Rollback Loops

The article analyzes the engineering challenges of moving configuration management from single-machine files to distributed cluster control planes, covering immutable versioning, schema validation layers, canary deployment by failure domains, two-phase commit for atomic switches, rollback with side-effect mitigation, observability across control and data planes, and maturity stages of config platforms.

Control Planecanary deploymentconfiguration management
0 likes · 47 min read
Safe Configuration Rollout at Scale: Versioning, Canary, and Rollback Loops
Data Bricklaying Diary
Data Bricklaying Diary
Sep 11, 2026 · Backend Development

Distributed Systems' Hardest Challenge: State Over Interfaces — Consistency, Idempotency & Async Tasks

This article explores why cross-service state management is harder than API design in distributed systems, using a refund example to illustrate how local transactions, idempotency, state machines, Outbox patterns, and manual intervention achieve recoverable convergence when operations span multiple services and external channels.

ConsistencyIdempotencyOutbox Pattern
0 likes · 21 min read
Distributed Systems' Hardest Challenge: State Over Interfaces — Consistency, Idempotency & Async Tasks
Architect's Guide
Architect's Guide
Sep 10, 2026 · Backend Development

Apollo Configuration Center: Complete Guide from Core Concepts to Kubernetes Deployment

This article provides a comprehensive tutorial on Apollo configuration center, covering its core concepts (application, environment, cluster, namespace), client architecture with long-polling and local caching, overall system design with Config/Admin services and Eureka, availability scenarios, hands-on SpringBoot integration with Maven setup, dynamic configuration testing (updates, rollbacks, offline fallback), multi-environment/cluster/namespace usage, and Kubernetes deployment via Docker and YAML manifests.

Configuration CenterCtripJava
0 likes · 27 min read
Apollo Configuration Center: Complete Guide from Core Concepts to Kubernetes Deployment
IT Services Circle
IT Services Circle
Sep 8, 2026 · Fundamentals

What Does await Actually Wait For? 10 Misconceptions That Break Async Code

This article dissects 10 common misconceptions about JavaScript's await keyword, explaining how it evaluates expressions, waits only for Promise settlement — not side effects — and why patterns like forEach with async callbacks, bare setTimeout, and Promise.all can cause silent bugs in production.

JavaScriptPromisesasync/await
0 likes · 16 min read
What Does await Actually Wait For? 10 Misconceptions That Break Async Code
AI Large-Model Wave and Transformation Guide
AI Large-Model Wave and Transformation Guide
Sep 8, 2026 · Artificial Intelligence

AI Agent Development: The Dual Challenge of Thinking Engineering & Distributed Systems

This article argues that AI agent development shifts from traditional coding to dual-system engineering: single agents require thinking logic design (prompt engineering, reasoning frameworks), while multi-agent systems demand distributed architecture skills (task graphs, state management, concurrency control), combining probabilistic reasoning with system reliability challenges.

AI agentsLLM AgentsLangGraph
0 likes · 14 min read
AI Agent Development: The Dual Challenge of Thinking Engineering & Distributed Systems
Cloud Architecture
Cloud Architecture
Sep 7, 2026 · Backend Development

Why a Single Timeout Spawned Two Risk Reviews: Production MCP Server Patterns

The article analyzes a timeout-induced duplicate risk review incident, then presents a comprehensive production-grade MCP server design covering stateless protocol alignment, schema validation, dual-key idempotency (request_key + business_key), UNKNOWN state machine, recovery workers, MCP Tasks integration, concurrency control, observability, and fault-injection testing to ensure exactly-once business effects.

IdempotencyMCPModel Context Protocol
0 likes · 27 min read
Why a Single Timeout Spawned Two Risk Reviews: Production MCP Server Patterns
Data Bricklaying Diary
Data Bricklaying Diary
Sep 7, 2026 · Backend Development

Independent Deployment ≠ Independent Evolution: Designing Dependencies, Routing & Release Architecture

The article explains why independent deployment doesn't guarantee independent service evolution, detailing five required capabilities—identifiable dependencies, compatible interfaces, controllable traffic routing, verifiable releases, and recoverable failures—using a refund service case study to illustrate dependency topology, interface evolution patterns, phased release strategies, and evidence-based rollback decisions.

Release Engineeringcanary deploymentdependency management
0 likes · 21 min read
Independent Deployment ≠ Independent Evolution: Designing Dependencies, Routing & Release Architecture
ITPUB
ITPUB
Sep 6, 2026 · Backend Development

How WeChat Resets 1 Billion Step Counts at Midnight Without Crashing

WeChat avoids server crashes during midnight step-count resets for 1 billion users by using logical time-based versioning instead of physical updates, a custom PaxosStore for atomic increments, delayed double-write buffers for clock skew, Redis ZSet sharding for rankings, and asynchronous cold-data archival during low-traffic hours.

High ConcurrencyPaxosStoreRedis
0 likes · 18 min read
How WeChat Resets 1 Billion Step Counts at Midnight Without Crashing
Spring Full-Stack Practical Cases
Spring Full-Stack Practical Cases
Sep 4, 2026 · Backend Development

10 Message Queue Patterns for Spring Boot: Decoupling, Async, Traffic Shaping & More

This article details 10 practical message queue scenarios for Spring Boot microservices, covering system decoupling, asynchronous processing, traffic shaping for flash sales, data synchronization, centralized logging, broadcast configuration updates, ordered message processing, delayed messages for timeouts, retry mechanisms with dead-letter queues, and transactional messaging for distributed consistency, with code examples for RabbitMQ, Kafka, and RocketMQ.

KafkaMessage QueueRabbitMQ
0 likes · 25 min read
10 Message Queue Patterns for Spring Boot: Decoupling, Async, Traffic Shaping & More
Programmer1970
Programmer1970
Aug 28, 2026 · Cloud Native

Nacos CP Mode Still Has Split-Brain? 4 Scenarios Where Raft Fails

This article explains why Nacos CP mode can still experience split-brain despite using Raft, detailing four trigger scenarios (even-node partitions, GC pauses, cross-AZ latency, snapshot failures), three reasons CP appears like AP (client cache, read paths, module separation), and five practical fixes plus diagnostic commands.

CP modeNacosRaft
0 likes · 13 min read
Nacos CP Mode Still Has Split-Brain? 4 Scenarios Where Raft Fails
dbaplus Community
dbaplus Community
Aug 27, 2026 · Backend Development

How MQ Message Reordering Triggered a Major Business Outage

A late‑night logistics alert revealed that out‑of‑order RocketMQ events caused order and shipment statuses to diverge, prompting a root‑cause analysis that uncovered a gift‑giving feature’s simultaneous event publishing and led to two mitigation strategies: delayed sending and ordered messages.

BackendJavaRocketMQ
0 likes · 8 min read
How MQ Message Reordering Triggered a Major Business Outage
Ray's Galactic Tech
Ray's Galactic Tech
Aug 25, 2026 · Backend Development

How to Secure Wallet Funds at Billion‑Scale Through Reconciliation: Balance Checks, Channel Matching, and Auto‑Repair

In high‑throughput wallet systems, reconciliation—covering internal balance validation, channel‑level matching, and controlled auto‑repair—acts as the final safeguard against fund discrepancies caused by lost callbacks, out‑of‑order events, duplicate entries, or concurrency conflicts, ensuring financial safety even with billions of daily transactions.

High Concurrencyauto-repairbalance-validation
0 likes · 33 min read
How to Secure Wallet Funds at Billion‑Scale Through Reconciliation: Balance Checks, Channel Matching, and Auto‑Repair
samdeepthink
samdeepthink
Aug 23, 2026 · Databases

Is Splitting Data Across Machines Distributed? Understand Sharding vs Distributed Systems

The article explains that merely placing data on multiple machines constitutes sharding—a way to split data for capacity—but true distributed systems require coordinated nodes that communicate, replicate, and handle failures, illustrated with an e‑commerce warehouse analogy and guidance on choosing between sharding and distributed databases.

database architecturedistributed systemsselection criteria
0 likes · 8 min read
Is Splitting Data Across Machines Distributed? Understand Sharding vs Distributed Systems
ITPUB
ITPUB
Aug 20, 2026 · Industry Insights

2026 China Database Technology Conference Launches: Data Fusion and AI Leadership

The 17th China Database Technology Conference (DTCC 2026) ran from August 20‑22 in Beijing, gathering top experts to discuss database kernel innovations, cloud‑native and distributed practices, AI‑driven data, vector databases, real‑time warehouses, and the emerging Agent era, while showcasing cutting‑edge solutions from Dameng, Tencent Cloud, Alibaba Cloud, OceanBase and GoldenDB.

AIAgentcloud native
0 likes · 15 min read
2026 China Database Technology Conference Launches: Data Fusion and AI Leadership
Xiaolin Talks Programming
Xiaolin Talks Programming
Aug 19, 2026 · Backend Development

Building Distributed Rate Limiting with Spring Boot, Redis & Lua: Sliding Window, Token Bucket & Flash Sale Defense

This article details a production-ready distributed rate limiting system using Spring Boot AOP, Redis Sorted Sets, and atomic Lua scripts, covering sliding window and token bucket algorithms, multi-dimensional flash sale protection, and benchmark results showing 100% accuracy with minimal latency overhead.

AOPLuaRedis
0 likes · 18 min read
Building Distributed Rate Limiting with Spring Boot, Redis & Lua: Sliding Window, Token Bucket & Flash Sale Defense
Mike Chen's Internet Architecture
Mike Chen's Internet Architecture
Aug 18, 2026 · Backend Development

How Alipay Ensures Payment Idempotency in High‑Concurrency Systems

The article explains why idempotency is critical for Alipay’s high‑concurrency, high‑reliability payment platform and outlines four key reasons—funds safety, handling network uncertainty, user experience, and distributed architecture—followed by concrete solutions such as unique business identifiers, database unique constraints, state‑machine control, and idempotent API design.

Database ConstraintsIdempotencybackend design
0 likes · 6 min read
How Alipay Ensures Payment Idempotency in High‑Concurrency Systems
Random Bulletin
Random Bulletin
Aug 18, 2026 · Operations

From Single‑Node to Distributed Monitoring Storage: Scaling TSDB for Million‑QPS

The article explains how a single‑machine Prometheus TSDB initially works well but eventually hits five scalability walls—capacity, single‑point failure, short retention, lack of global view, and throughput limits—and then details the step‑by‑step evolution to remote_write with object‑storage‑backed Thanos and finally to native distributed TSDBs such as Cortex, Mimir, and VictoriaMetrics, including their trade‑offs, costs, and practical selection guidance.

CortexPrometheusTSDB
0 likes · 21 min read
From Single‑Node to Distributed Monitoring Storage: Scaling TSDB for Million‑QPS
Yumin Fish Harvest
Yumin Fish Harvest
Aug 18, 2026 · Backend Development

Distributed Snowflake ID Generation in Java: Hutool and MyBatis-Plus Demo

This article walks through the Snowflake algorithm for distributed ID generation, detailing its timestamp, workerId and sequence fields, handling clock rollback and sequence overflow, and provides a complete Java implementation with integration examples for Hutool and MyBatis‑Plus, including configuration and test code.

HutoolJavaMyBatis-Plus
0 likes · 28 min read
Distributed Snowflake ID Generation in Java: Hutool and MyBatis-Plus Demo
Machine Heart
Machine Heart
Aug 17, 2026 · Artificial Intelligence

TensorCast Cuts First‑Token Latency by Up to 93.2% with Unified Programmable Tensor Management

TensorCast introduces a unified, programmable tensor lifecycle layer for large‑model infrastructure, achieving up to a 93.2% reduction in first‑token latency, a 228.6× speed‑up in model startup, and performance comparable to specialized KV‑cache systems while simplifying development.

LLM infrastructureTensorCastdistributed systems
0 likes · 11 min read
TensorCast Cuts First‑Token Latency by Up to 93.2% with Unified Programmable Tensor Management
Cloud Architecture
Cloud Architecture
Aug 16, 2026 · Backend Development

How to Build a 100k QPS Seckill System with Spring Boot, Redis, and Lua

This article provides a production‑grade, step‑by‑step engineering guide for designing a high‑concurrency seckill (flash‑sale) system that can sustain 100,000 QPS using Spring Boot, Redis with Lua scripts, asynchronous messaging, and comprehensive fault‑tolerance, monitoring, and scalability techniques.

High ConcurrencyLuaRedis
0 likes · 40 min read
How to Build a 100k QPS Seckill System with Spring Boot, Redis, and Lua
Xike
Xike
Aug 16, 2026 · Backend Development

Idempotency Strategies for APIs: Preventing Duplicate Submissions

The article analyzes why network timeouts, client retries, user double‑clicks, and at‑least‑once MQ delivery cause duplicate intents, then details five practical idempotency techniques—conditional updates, unique constraints, Idempotency‑Key tokens, state machines with optimistic locks, and distributed locks—along with their trade‑offs and monitoring tips.

APIIdempotencyToken
0 likes · 11 min read
Idempotency Strategies for APIs: Preventing Duplicate Submissions
21CTO
21CTO
Aug 15, 2026 · Cloud Native

Ryan Dahl Unveils celld – A Self‑Hosted Distributed Durable Objects Platform

Ryan Dahl, the creator of Node.js, announced celld—a self‑hosted, distributed implementation of Cloudflare Workers and Durable Objects that promises lower costs, uses S3 storage, Rust/Tokio runtime, and supports stateful serverless workloads such as real‑time collaboration and AI agents.

CloudflareDurable ObjectsRust
0 likes · 6 min read
Ryan Dahl Unveils celld – A Self‑Hosted Distributed Durable Objects Platform
Ray's Galactic Tech
Ray's Galactic Tech
Aug 14, 2026 · Backend Development

Payment Callback After Order Cancellation? 3 Consistency Challenges and Practical Engineering Solutions

The article examines why payment callbacks that arrive later than order cancellations cause money‑in‑order‑canceled anomalies, outlines the three core eventual‑consistency problems—message ordering, duplicate consumption, and message loss—and presents a complete engineering solution using business idempotency keys, a transactional outbox, state‑machine handling, and reconciliation compensation.

IdempotencyKafkadistributed systems
0 likes · 35 min read
Payment Callback After Order Cancellation? 3 Consistency Challenges and Practical Engineering Solutions
DataFunSummit
DataFunSummit
Aug 13, 2026 · Cloud Native

Agent Architecture Evolution: From Monolithic Self‑Management to Distributed Hosting

The article outlines a step‑by‑step evolution of Agent systems, explaining why traditional microservice patterns fail, describing three monolithic deployment models, detailing how separating session, memory, and environment state enables distributed hosting, and presenting function‑as‑a‑service to fully managed ReAct and multi‑Agent collaboration via Registry and A2A.

AgentAgent RegistryFunction-as-a-Service
0 likes · 14 min read
Agent Architecture Evolution: From Monolithic Self‑Management to Distributed Hosting
AI Engineering
AI Engineering
Aug 13, 2026 · Artificial Intelligence

When AI Agents Become Internet Users: Exploring the AI‑SNS Network

The article examines how the internet, traditionally built for human users, may evolve into an AI‑agent‑centric network, discussing the need for agent discovery, social relationships, and a new infrastructure illustrated by the open‑source AI‑SNS project.

AI InfrastructureAI agentsAI‑SNS
0 likes · 11 min read
When AI Agents Become Internet Users: Exploring the AI‑SNS Network
ITPUB
ITPUB
Aug 12, 2026 · Backend Development

How Dingdang Kuaike Achieves 28‑Minute Drug Delivery: The Underlying Tech Architecture

In this interview, Dingdang Kuaike’s R&D head Song Zilong explains how the company transformed a tightly coupled monolith into a micro‑service system, built intelligent real‑time scheduling, migrated Oracle to sharded MySQL across hybrid‑cloud IDC, and leveraged AI to sustain a 28‑minute drug‑delivery promise.

AI integrationR&D Managementdatabase sharding
0 likes · 17 min read
How Dingdang Kuaike Achieves 28‑Minute Drug Delivery: The Underlying Tech Architecture
Woodpecker Software Testing
Woodpecker Software Testing
Aug 12, 2026 · Cloud Native

Distributed vs Monolithic: Uncovering the Real Performance Trade‑offs

The article analyzes distributed and traditional monolithic architectures across response latency, throughput, scalability, and fault‑tolerance cost, revealing that distributed systems introduce network latency, serialization overhead, higher resource consumption, and complex failure modes that often offset their scalability benefits, and provides concrete case studies from Netflix, Alibaba, and LinkedIn.

Throughputcloud nativedistributed systems
0 likes · 7 min read
Distributed vs Monolithic: Uncovering the Real Performance Trade‑offs
Coder Life Journal
Coder Life Journal
Aug 11, 2026 · Backend Development

Spring Cloud + Kafka: 6 Common Pitfalls and How to Avoid Them

The article walks through six real‑world pitfalls when integrating Spring Cloud with Kafka—message loss, duplicate processing, out‑of‑order events, massive consumer lag, serialization mismatches, and misuse of Kafka transactions—and provides concrete configuration tweaks, code examples, and operational safeguards to prevent each issue.

KafkaSpring Clouddistributed systems
0 likes · 9 min read
Spring Cloud + Kafka: 6 Common Pitfalls and How to Avoid Them
Thought Artisan
Thought Artisan
Aug 8, 2026 · Fundamentals

Why Software Architecture Is Hard: 8 Core Trade-offs from 'The Hard Parts'

This article reviews 'Software Architecture: The Hard Parts', highlighting eight key architectural challenges including trade-offs, data-architecture tensions, service granularity, database decomposition, distributed data access, contract design, code reuse, and error handling in distributed systems.

contract designdatabase decompositiondistributed systems
0 likes · 7 min read
Why Software Architecture Is Hard: 8 Core Trade-offs from 'The Hard Parts'
Cloud Architecture
Cloud Architecture
Aug 7, 2026 · Backend Development

Designing an Industrial‑Grade Message Queue for Tens of Millions of Orders

This article presents a step‑by‑step design of HermesMQ, an industrial‑grade message queue built from scratch to support ten‑million‑order traffic, covering storage as sequential logs, network architecture with Netty and Reactor, high‑availability replication, partition ordering, transaction messaging, back‑pressure, observability, and practical deployment guidelines.

JavaMessage QueueTransaction Messaging
0 likes · 45 min read
Designing an Industrial‑Grade Message Queue for Tens of Millions of Orders
AntData
AntData
Aug 7, 2026 · Big Data

Designing Lampara: Ant Group’s Real‑Time Data Processing System for End‑to‑End SLA Guarantees

The article details Ant Group’s Lampara, a next‑generation real‑time data processing platform that embeds end‑to‑end SLA guarantees, active disaster‑recovery, scenario‑driven development and enhanced operators, showing how it improves reliability, efficiency and cost while supporting large‑scale business scenarios such as flash sales, AI agents and marketing campaigns.

LamparaSLAactive disaster recovery
0 likes · 18 min read
Designing Lampara: Ant Group’s Real‑Time Data Processing System for End‑to‑End SLA Guarantees
Node.js Tech Stack
Node.js Tech Stack
Aug 7, 2026 · Backend Development

Why Ryan Dahl Defied His Promise and Built the celld JavaScript Runtime

After vowing never to create another JavaScript runtime, Ryan Dahl spent over a year developing celld, an open‑source, self‑hosted distributed runtime that brings Cloudflare Workers and Durable Objects to your own machines, using V8, SQLite, Rust async, and S3 for coordination, while exposing its performance trade‑offs and early‑stage limitations.

Cloudflare WorkersDurable ObjectsJavaScript runtime
0 likes · 12 min read
Why Ryan Dahl Defied His Promise and Built the celld JavaScript Runtime
Xike
Xike
Aug 5, 2026 · Backend Development

How to Automatically Cancel Unpaid Orders When They Timeout

The article explains a reliable, idempotent solution for automatically cancelling orders that remain unpaid after a configured deadline, covering data modeling, state transitions, trigger mechanisms using delayed messages or scans, handling race conditions with payment, and essential monitoring and pitfalls.

BackendMessage QueueRocketMQ
0 likes · 17 min read
How to Automatically Cancel Unpaid Orders When They Timeout
AI Open-Source Efficiency Guide
AI Open-Source Efficiency Guide
Aug 5, 2026 · Artificial Intelligence

Orchard: Microsoft’s Open‑Source Agent Framework Hits 0.28 s Latency with 1,000 Sandboxes

Orchard is Microsoft’s open‑source, Kubernetes‑native agent modeling platform that isolates execution in lightweight sandboxes, separates control‑plane operations, supports arbitrary base images and multiple built‑in harnesses, and—according to official benchmarks—delivers an average command latency of 0.28 seconds when running 1,000 concurrent sandboxes.

AI agentsKubernetesOrchard
0 likes · 17 min read
Orchard: Microsoft’s Open‑Source Agent Framework Hits 0.28 s Latency with 1,000 Sandboxes
Data Party THU
Data Party THU
Aug 4, 2026 · Operations

Why Multi-Agent Systems Are Fundamentally Distributed Systems

The article argues that multi‑agent workflows behave like traditional distributed systems, showing how deadlocks, state pollution, and silent drift arise from coordination failures rather than AI shortcomings, and it offers concrete engineering practices—timeouts, idempotency, cycle detection, and audit trails—to build reliable production‑grade agent pipelines.

deadlockdistributed systemsmulti-agent systems
0 likes · 14 min read
Why Multi-Agent Systems Are Fundamentally Distributed Systems
DeepNoMind
DeepNoMind
Aug 2, 2026 · Databases

Understand Partitioning vs Sharding in 5 Minutes

The article explains how partitioning splits tables within a single database and how sharding distributes data across multiple database instances, comparing their types, advantages, limitations, and trade‑offs, and provides practical examples and a decision framework for choosing the right strategy.

Partitioningdatabasesdistributed systems
0 likes · 7 min read
Understand Partitioning vs Sharding in 5 Minutes
IT Learning Made Simple
IT Learning Made Simple
Aug 1, 2026 · Backend Development

How Military Command Structures Reveal the Secrets of Large‑Scale System Architecture

The article draws a detailed analogy between the People's Army command hierarchy and modern distributed system design, mapping each military layer to software architecture components and highlighting fault tolerance, unified standards, elastic scaling, and comprehensive security as lessons for IT professionals.

Elastic ScalingSecurity Architecturedistributed systems
0 likes · 9 min read
How Military Command Structures Reveal the Secrets of Large‑Scale System Architecture
DeepHub IMBA
DeepHub IMBA
Jul 28, 2026 · Artificial Intelligence

Why Multi‑Agent Systems Are Fundamentally Distributed Systems

Multi‑agent workflows often deadlock or drift because their agents behave like distributed nodes, so treating them as a distributed system reveals classic failure modes—deadlocks, state pollution, lack of timeouts, and missing idempotency—allowing proven engineering practices to keep AI pipelines reliable.

AI engineeringLangGraphdistributed systems
0 likes · 14 min read
Why Multi‑Agent Systems Are Fundamentally Distributed Systems
IT Learning Made Simple
IT Learning Made Simple
Jul 27, 2026 · R&D Management

Six-Month Study Plan for System Architecture Designer Exam: Build Foundations, Aim for 70+ Scores

This guide outlines a detailed 24‑week, 360‑hour preparation roadmap for the System Architecture Designer certification, targeting working professionals and beginners, dividing the study into six phases—from entry to adjustment—each with specific weekly tasks, learning topics, practice exams, and milestones to achieve a 70+ score.

Study Plancertificationdistributed systems
0 likes · 13 min read
Six-Month Study Plan for System Architecture Designer Exam: Build Foundations, Aim for 70+ Scores
dbaplus Community
dbaplus Community
Jul 26, 2026 · Cloud Native

Will AI Replace Kubernetes? Co‑Founder Brendan Burns on Its Rise and End

Brendan Burns recounts how he convinced Google to back Kubernetes, built the MVP in five days, navigated open‑source governance, tackled technical challenges like Etcd and declarative design, expanded the platform for AI workloads, and reflects on why even successful software like Kubernetes inevitably faces obsolescence.

AI workloadsKubernetescloud native
0 likes · 31 min read
Will AI Replace Kubernetes? Co‑Founder Brendan Burns on Its Rise and End
Cloud Architecture
Cloud Architecture
Jul 26, 2026 · Backend Development

Seckill System Architecture: 7 Core Design Strategies for High-Concurrency Sales

This article presents a comprehensive, step‑by‑step analysis of building a flash‑sale (seckill) system that can survive instant traffic spikes, detailing seven essential design ideas such as static page delivery, token gating, Redis atomic decrement, asynchronous queuing, multi‑layer rate limiting, service isolation, idempotent processing, and end‑to‑end monitoring and recovery.

High ConcurrencyIdempotencyLua
0 likes · 28 min read
Seckill System Architecture: 7 Core Design Strategies for High-Concurrency Sales
Dabaoshi
Dabaoshi
Jul 26, 2026 · Databases

When to Scale MySQL: From Single Instance to Distributed Architecture

The article explains why a single MySQL server eventually hits read, write, or availability limits, outlines a step‑by‑step evolution—from SQL tuning and caching to read‑write splitting, high‑availability setups, and finally vertical or horizontal sharding—while warning against premature distribution.

MySQLRead-Write Splittingdistributed systems
0 likes · 15 min read
When to Scale MySQL: From Single Instance to Distributed Architecture
Cloud Architecture
Cloud Architecture
Jul 25, 2026 · Backend Development

From MySQL to 10M QPS: A Full‑Stack Engineering Blueprint for High‑Throughput Transaction Systems

This white‑paper dissects why a monolithic MySQL‑based order service collapses under peak traffic and presents a layered, asynchronous architecture—using Redis for stock pre‑allocation, RocketMQ for transactional messaging, sharding, idempotency, and comprehensive observability—to reliably handle tens of millions of queries per second.

MySQLRedisRocketMQ
0 likes · 34 min read
From MySQL to 10M QPS: A Full‑Stack Engineering Blueprint for High‑Throughput Transaction Systems
Geek Labs
Geek Labs
Jul 24, 2026 · Artificial Intelligence

How Google’s New Open‑Source Projects Make AI Agents Production‑Ready

Google Cloud recently open‑sourced two Go projects—Scion, which isolates and coordinates multiple AI agents, and AX, a distributed runtime that enables a single long‑running agent to resume after failures—detailing their architectures, usage steps, real‑world use cases, limitations, and the broader strategy of turning agents from experimental toys into reliable production workers.

AI agentsAXGo
0 likes · 12 min read
How Google’s New Open‑Source Projects Make AI Agents Production‑Ready
IT Learning Made Simple
IT Learning Made Simple
Jul 23, 2026 · Fundamentals

50 Essential Architecture Concepts in One Sentence Each

This article presents 50 concise statements that cover core software architecture concepts, including fundamentals, types, views, design principles, practical practices, distributed architecture basics, and the mindset needed for architects, providing a quick reference for beginners.

architect mindsetarchitecture fundamentalsdesign principles
0 likes · 10 min read
50 Essential Architecture Concepts in One Sentence Each
Top Architect
Top Architect
Jul 23, 2026 · Backend Development

How Taobao’s Backend Architecture Evolved Over a Decade

The article walks through Taobao’s backend architecture transformation from a single‑server setup to a cloud‑native, micro‑service ecosystem, detailing fourteen evolutionary stages—including separate Tomcat and DB, caching, load balancing, sharding, NoSQL, ESB, containerization, and cloud deployment—while highlighting key concepts, challenges, and design principles.

Cachingbackend architecturecloud native
0 likes · 23 min read
How Taobao’s Backend Architecture Evolved Over a Decade
IT Learning Made Simple
IT Learning Made Simple
Jul 22, 2026 · Fundamentals

Architect’s Reading List: From Beginner to Master

This article presents a curated reading list for software architects, organized by career stages and covering design patterns, code quality, architecture, distributed systems, cloud‑native topics, along with reading principles, recommendations, and a top‑10 book ranking to guide continuous learning.

Design Patternsbook listcloud native
0 likes · 10 min read
Architect’s Reading List: From Beginner to Master
IT Learning Made Simple
IT Learning Made Simple
Jul 21, 2026 · Fundamentals

Key Takeaways from 'Designing Large-Scale Distributed Systems'

This note distills the core engineering practices for building and operating large‑scale distributed systems, covering system definition, distributed vs single‑node trade‑offs, CAP theorem choices, consistency levels, transaction patterns, load‑balancing algorithms, cache strategies, message‑queue reliability, coordination services like ZooKeeper, and essential design principles.

CAP theoremCachingConsistency
0 likes · 11 min read
Key Takeaways from 'Designing Large-Scale Distributed Systems'
Golang Shines
Golang Shines
Jul 20, 2026 · Cloud Native

7 Golden Rules for Building High‑Availability Cloud‑Native Go Services (Production‑Proven)

This article presents a step‑by‑step guide to building highly available cloud‑native Go systems, covering graceful error handling, structured logging, minimal dependencies, concurrency control, health checks, Raft‑based replication, timeout/retry strategies, circuit breaking, rate limiting, observability with Zap, Loki, Prometheus, OpenTelemetry, and future architectural directions.

Gocloud nativedistributed systems
0 likes · 18 min read
7 Golden Rules for Building High‑Availability Cloud‑Native Go Services (Production‑Proven)
Data Party THU
Data Party THU
Jul 18, 2026 · Artificial Intelligence

Smart Cellular Bricks: 3D Neural Cellular Automata for Life‑Like Modular Robots

The study introduces Smart Cellular Bricks, a modular robot system that uses 3D Neural Cellular Automata to enable identical cubes to exchange minimal local information, achieve 98.97% shape‑classification accuracy, 94.8% damage‑detection precision, and self‑repair within 60 update cycles, demonstrating scalable, life‑like collective intelligence.

Sakana AIdistributed systemsmodular robotics
0 likes · 7 min read
Smart Cellular Bricks: 3D Neural Cellular Automata for Life‑Like Modular Robots
Random Bulletin
Random Bulletin
Jul 17, 2026 · Backend Development

Ten‑Million QPS Architecture: From Strong to Eventual Consistency with Layered Design

The article examines why ultra‑high‑throughput systems must move from costly strong consistency to eventual consistency, explaining consistency layering, the trade‑offs of coordination latency, availability, and tail latency, and how local message tables, idempotent delivery, compensation (Saga) and periodic reconciliation together ensure data eventually aligns without sacrificing performance.

ConsistencySAGAcompensation
0 likes · 17 min read
Ten‑Million QPS Architecture: From Strong to Eventual Consistency with Layered Design
Ray's Galactic Tech
Ray's Galactic Tech
Jul 17, 2026 · Backend Development

Three Critical Guarantees for Message Queues: No Loss, No Duplicates, No Disorder – Deep Dive into Production‑Grade Solutions

This article dissects why modern systems must enforce three reliability guarantees—no message loss, no duplicate processing, and no out‑of‑order delivery—by examining real‑world order flows, outbox patterns, idempotent keys, partitioning strategies, consumer acknowledgments, and operational safeguards such as replay, dead‑letter handling, and monitoring.

IdempotencyKafkaMessage Queue
0 likes · 28 min read
Three Critical Guarantees for Message Queues: No Loss, No Duplicates, No Disorder – Deep Dive into Production‑Grade Solutions
Cloud Architecture
Cloud Architecture
Jul 16, 2026 · Backend Development

Beyond Delayed Double Delete: How CDC Closed‑Loop Governance Solves Cache Consistency

The article dissects why the classic "update‑DB‑then‑delete‑cache" or its reverse is only a probability fix for cache inconsistency, demonstrates the failure modes of delayed double delete under high load and replication lag, and presents a production‑grade solution built on Binlog CDC with versioning, four‑plane governance, and robust error handling to achieve reliable cache synchronization.

BackendCDCConsistency
0 likes · 36 min read
Beyond Delayed Double Delete: How CDC Closed‑Loop Governance Solves Cache Consistency
Coder Trainee
Coder Trainee
Jul 16, 2026 · Interview Experience

Java Interview Essentials: 10 Must‑Ask Microservice Questions

This article presents ten essential microservice interview questions, covering service registration and discovery, API gateways, fault tolerance, distributed transactions, tracing, configuration management, messaging, ID generation, and service decomposition, each explained with principles, diagrams, code snippets, and tips for impressing interviewers.

BackendJavaSpring
0 likes · 14 min read
Java Interview Essentials: 10 Must‑Ask Microservice Questions
Ray's Galactic Tech
Ray's Galactic Tech
Jul 15, 2026 · Artificial Intelligence

Scalable Knowledge Base with High‑Concurrency Crawling and Vector Search

The article explains why a production‑grade enterprise knowledge base requires more than just dumping PDFs into a vector store, detailing a distributed, event‑driven architecture with separate collection, processing, retrieval, and governance layers that handle high‑concurrency crawling, real‑time cleaning, versioned indexing, permission filtering, and feedback‑driven updates.

Knowledge BaseRAGdata pipeline
0 likes · 39 min read
Scalable Knowledge Base with High‑Concurrency Crawling and Vector Search
IT Services Circle
IT Services Circle
Jul 14, 2026 · Backend Development

Designing a Restaurant Reservation System: How to Slice Time Granularity?

The article compares interview‑style answers and real‑world production for a restaurant reservation system, outlining six design dimensions, three time‑slot strategies, five engineering challenges, and architectural considerations to help engineers choose the right granularity and avoid common pitfalls.

backend designconcurrency controldistributed systems
0 likes · 15 min read
Designing a Restaurant Reservation System: How to Slice Time Granularity?
Shuge Unlimited
Shuge Unlimited
Jul 12, 2026 · Databases

Milvus 3.0 Streaming Architecture: 16 PChannels, Five Interceptor Layers, and a Self‑Built WAL 5.8× Faster Than Kafka

Milvus 3.0 replaces the dual‑track write path of 2.x with a unified WAL‑first design, introduces a three‑layer channel model (PChannel, VChannel, CChannel) and a five‑layer interceptor chain, adds the Woodpecker WAL that outperforms Kafka/Pulsar by up to 5.8×, and provides pluggable back‑ends, atomic broadcasting, and provable recovery mechanisms.

MilvusStreaming ArchitectureWAL
0 likes · 27 min read
Milvus 3.0 Streaming Architecture: 16 PChannels, Five Interceptor Layers, and a Self‑Built WAL 5.8× Faster Than Kafka
Linyb Geek Road
Linyb Geek Road
Jul 12, 2026 · Operations

Designing a High‑Availability Architecture: Core Principles and Practices

This article outlines the essential principles for building a high‑availability system, covering cluster and distributed designs, fault‑tolerance, reliable hardware, disaster recovery, monitoring, security, capacity planning, and automated scaling to achieve optimal performance and resilience.

automationcapacity planningcluster architecture
0 likes · 6 min read
Designing a High‑Availability Architecture: Core Principles and Practices
Random Bulletin
Random Bulletin
Jul 11, 2026 · Backend Development

Why Fixed Timeouts Cause Snowball Failures and How Tiered Timeouts Save a 10M QPS System

A 3‑second fixed timeout turned an 800 ms latency spike in a recommendation service into a full‑site outage, illustrating how static timeouts can exhaust thread pools; the article walks through evolving from global timeouts to per‑interface limits, deadline propagation, and coordinated timeout‑retry‑circuit‑breaker strategies for resilient 10 M‑QPS systems.

RPCRetrycircuit breaker
0 likes · 15 min read
Why Fixed Timeouts Cause Snowball Failures and How Tiered Timeouts Save a 10M QPS System
YiSu Grain
YiSu Grain
Jul 6, 2026 · Fundamentals

Understanding CAP and BASE Through a Simple Network Partition Example

The article explains the CAP theorem and BASE model by walking through a concrete scenario of two data centers losing network connectivity, showing how architects must choose between consistency and availability and illustrating typical CP and AP use cases.

BASECAP theoremConsistency
0 likes · 9 min read
Understanding CAP and BASE Through a Simple Network Partition Example
Subtle Storm
Subtle Storm
Jul 1, 2026 · Backend Development

How to Tackle the “Three Highs” of Internet Systems Without Burning Out

The article analyzes the intertwined challenges of high concurrency, high performance, and high availability in internet services, explains why they cannot all be maximized simultaneously, and presents concrete architectural tactics—partitioning, caching, async processing, redundancy, and CAP trade‑offs—to achieve a balanced, resilient system.

CAP theoremCachingHigh Concurrency
0 likes · 7 min read
How to Tackle the “Three Highs” of Internet Systems Without Burning Out
Random Bulletin
Random Bulletin
Jul 1, 2026 · Operations

Evolving Message Expiration for 10 Million QPS: From No TTL to Smart Policies

When a high‑traffic system processes billions of messages per day, stale “zombie” messages can corrupt business state; this article walks through the evolution from never‑expiring queues to uniform TTL, then per‑topic and per‑message TTL, and finally to smart, context‑aware expiration, detailing the engineering trade‑offs, implementation patterns, and operational checklist needed for reliable 10 M‑QPS message pipelines.

KafkaMessage QueueRocketMQ
0 likes · 19 min read
Evolving Message Expiration for 10 Million QPS: From No TTL to Smart Policies
FunTester
FunTester
Jul 1, 2026 · Operations

When One Timeout Triggers a Platform‑Wide Outage

The article explains how unbounded retries, replication fan‑out, and naïve autoscaling can amplify a single timeout into a cascade of failures, and it proposes bounded retry policies, load‑aware scaling, and layered persistence as safeguards for reliable API‑centric systems.

AutoscalingReplicationRetry
0 likes · 12 min read
When One Timeout Triggers a Platform‑Wide Outage
BanTech Think Tank
BanTech Think Tank
Jun 30, 2026 · Industry Insights

Evolution of Banking Core System Architecture: Historical Review and Future Trends

This article examines the three major phases of Chinese commercial banks' core system architecture—system budding, centralized mainframe, and distributed designs—analyzes the technical improvements and shortcomings of each generation, and forecasts post‑distributed trends such as domestic‑technology adoption, cloud‑native deployment, intelligent operations, real‑time data warehousing, micro‑service structures, and large‑model integration.

Large Language ModelsReal-time Data Warehousearchitecture evolution
0 likes · 23 min read
Evolution of Banking Core System Architecture: Historical Review and Future Trends
Random Bulletin
Random Bulletin
Jun 26, 2026 · Operations

Message Duplication Is Inevitable: Building a Multi‑Layer Idempotency Middleware for Ten‑Million QPS

Message queues guarantee at‑least‑once delivery, making duplicate messages a normal feature; the article examines a real coupon‑distribution incident, critiques business‑level idempotency approaches, and outlines a layered platform‑wide middleware design—including unique keys, state machines, storage choices, and TTL strategies—to achieve reliable processing at ten‑million QPS scale.

IdempotencyMessage Queuedistributed systems
0 likes · 19 min read
Message Duplication Is Inevitable: Building a Multi‑Layer Idempotency Middleware for Ten‑Million QPS
Random Bulletin
Random Bulletin
Jun 25, 2026 · Operations

Scaling Message Queues to 10M QPS: From Downtime to Seamless Online Expansion

At the 10‑million‑QPS scale, expanding a message‑queue cluster no longer hinges on simply adding brokers; it requires coordinated online upgrades of metadata, data migration with dynamic throttling, cooperative consumer rebalance, shadow‑traffic warm‑up, and rollback snapshots, making the act of adding machines the hardest part.

10M QPSMessage QueueRebalance
0 likes · 27 min read
Scaling Message Queues to 10M QPS: From Downtime to Seamless Online Expansion
Cloud Architecture
Cloud Architecture
Jun 24, 2026 · Backend Development

Four Levels of Concurrency Control: Optimistic Locks to Message Queues for Million‑QPS

High‑concurrency systems must go beyond simple locking; the article breaks down four progressive strategies—optimistic locking, database pessimistic locking, Redis distributed locks, and finally message‑queue serialization—explaining their trade‑offs, implementation details, pitfalls, and how to combine them into a layered architecture that sustains million‑QPS workloads with consistency, throughput, and recoverability.

Message QueueOptimistic LockPessimistic Lock
0 likes · 32 min read
Four Levels of Concurrency Control: Optimistic Locks to Message Queues for Million‑QPS
Random Bulletin
Random Bulletin
Jun 24, 2026 · Operations

Rack Awareness: From Zero to High‑QPS – Boost Availability, Cut Cross‑AZ Traffic

A real‑world rack‑power outage showed that three‑replica Kafka clusters can lose all replicas when brokers share a failure domain, prompting a deep dive into rack awareness—how fault‑domain tags are injected, replica‑placement and leader‑distribution algorithms, consumer‑proximity reads, bandwidth costs, failure scenarios, and the stepwise evolution from hundred‑thousand to ten‑million QPS deployments.

KafkaPulsarRack Awareness
0 likes · 26 min read
Rack Awareness: From Zero to High‑QPS – Boost Availability, Cut Cross‑AZ Traffic
Random Bulletin
Random Bulletin
Jun 23, 2026 · Backend Development

Choosing the Right Replication Factor for 10M QPS Message Queues: From Dual to Multi‑Replica

The article walks through why a single replica is insufficient for million‑scale message queues, explains the latency and data‑loss trade‑offs of dual‑replica sync and async modes, shows how three‑replica majority voting becomes the sweet spot, and then details the engineering considerations for scaling to five or more replicas across availability zones and regions.

AZ awarenessKafkaMessage Queue
0 likes · 18 min read
Choosing the Right Replication Factor for 10M QPS Message Queues: From Dual to Multi‑Replica