Tagged articles

high QPS

43 articles · Page 1 of 1
Random Bulletin
Random Bulletin
Aug 26, 2026 · Operations

Log Standards at Million‑QPS Scale: From Free‑form to Strict Structured Logging

When a production outage forces a midnight investigation across five services, the lack of log standards turns a quick debug into an all‑night forensic hunt; the article explains how structured logs, a unified schema, level semantics, trace_id linking, and field‑level masking enforced by SDK, Lint and CI can eliminate these five walls and make logging scalable, searchable, and compliant.

CI enforcementStructured Logginghigh QPS
0 likes · 22 min read
Log Standards at Million‑QPS Scale: From Free‑form to Strict Structured Logging
Random Bulletin
Random Bulletin
Aug 24, 2026 · Operations

Scaling Log Volumes from GB to TB at Ten‑Million QPS: Cost‑Effective Strategies and Architecture

At ten‑million QPS, log data can explode from a few gigabytes to terabytes or even petabytes, triggering storage blow‑up, pipeline saturation, slow queries, runaway costs, and poor signal‑to‑noise, and the article breaks down ingest, index, and store costs while presenting edge sampling, label‑based indexing, tiered storage, and log‑to‑metric rollup as mitigation tactics.

Cost OptimizationElasticsearchLoki
0 likes · 19 min read
Scaling Log Volumes from GB to TB at Ten‑Million QPS: Cost‑Effective Strategies and Architecture
Random Bulletin
Random Bulletin
Aug 7, 2026 · Operations

From Simple Toggles to Platform‑Scale Switch Management for Million‑QPS Systems

The article explains how moving feature switches from hard‑coded if‑else statements to a platform‑based runtime system enables granular targeting, gradual rollouts, A/B experiments, instant kill‑switches, and SDK‑based local evaluation that can handle millions of queries per second without redeployments.

A/B testingFeature FlagsSDK local evaluation
0 likes · 17 min read
From Simple Toggles to Platform‑Scale Switch Management for Million‑QPS Systems
Random Bulletin
Random Bulletin
Jul 21, 2026 · Backend Development

Thread Pool Design for Million‑QPS Systems: From Shared to Isolated

The article explains how a shared thread pool can turn a localized slowdown into a system‑wide outage, then walks through the bulkhead isolation pattern, bounded queues, trade‑offs between thread‑pool and semaphore isolation, sizing formulas, monitoring metrics, and dynamic tuning for high‑QPS services.

ConcurrencyThread Poolbulkhead pattern
0 likes · 16 min read
Thread Pool Design for Million‑QPS Systems: From Shared to Isolated
Cloud Architecture
Cloud Architecture
Jul 20, 2026 · Databases

Designing MySQL for Millions of QPS: From Single Server to Distributed Architecture

The article walks through a real‑world order system that spikes to 300,000 QPS, explaining why the original single‑node MySQL design fails, and detailing a step‑by‑step evolution—index tuning, transaction fixes, read‑write splitting, vertical and horizontal sharding, plus data‑pipeline integration—to achieve stable low latency at massive scale.

Database ScalingIndex OptimizationMySQL
0 likes · 20 min read
Designing MySQL for Millions of QPS: From Single Server to Distributed Architecture
Random Bulletin
Random Bulletin
Jul 17, 2026 · Backend Development

Why Moving from Real‑Time to Batch Is Essential for Scaling to Tens of Millions QPS

Scaling a service from millions to tens of millions of queries per second fails not because of data size but due to per‑request fixed costs, and the article shows how batching aggregates these costs, dramatically boosts throughput, reduces latency, and introduces new challenges such as memory pressure and partial failures.

Batch ProcessingSystem Designhigh QPS
0 likes · 17 min read
Why Moving from Real‑Time to Batch Is Essential for Scaling to Tens of Millions QPS
Random Bulletin
Random Bulletin
Jul 16, 2026 · Backend Development

Task Queues at Ten‑Million QPS: From Synchronous Processing to Queue‑Based Architecture

The article explains how moving non‑critical actions in a high‑traffic e‑commerce order flow to a task queue separates core work from optional work, solving delivery guarantees, idempotency, backlog management, and ordering challenges while dramatically reducing latency and fault contagion.

Task Queueasynchronous-processingbacklog management
0 likes · 17 min read
Task Queues at Ten‑Million QPS: From Synchronous Processing to Queue‑Based Architecture
Random Bulletin
Random Bulletin
Jul 13, 2026 · Backend Development

Why Simple Retry Triggers Avalanches and How Smart Strategies Save 10M‑QPS Systems

The article explains how a naïve ‘retry three times’ default can amplify a minor 200 ms latency spike into a full‑scale outage in a 10 million‑QPS system, and walks through a progressive design—from identifying retry‑eligible transient faults, adding exponential backoff with jitter, enforcing idempotency, applying token‑bucket retry budgets, integrating circuit breakers, to using hedged requests—showing how each layer prevents traffic amplification and ensures reliable high‑throughput services.

Circuit BreakerJitterbackoff
0 likes · 18 min read
Why Simple Retry Triggers Avalanches and How Smart Strategies Save 10M‑QPS Systems
Random Bulletin
Random Bulletin
Jul 11, 2026 · Backend Development

Why Fixed Timeouts Cause Snowball Failures and How Tiered Timeouts Save a 10M QPS System

A 3‑second fixed timeout turned an 800 ms latency spike in a recommendation service into a full‑site outage, illustrating how static timeouts can exhaust thread pools; the article walks through evolving from global timeouts to per‑interface limits, deadline propagation, and coordinated timeout‑retry‑circuit‑breaker strategies for resilient 10 M‑QPS systems.

Circuit BreakerRPCRetry
0 likes · 15 min read
Why Fixed Timeouts Cause Snowball Failures and How Tiered Timeouts Save a 10M QPS System
Random Bulletin
Random Bulletin
Jul 10, 2026 · Backend Development

Why High‑QPS Systems Move from JSON to Efficient Binary Serialization

A flame‑graph reveals that in high‑throughput services only about 30% of CPU time is spent on business logic while the rest is consumed by serialization, prompting a step‑by‑step evolution from flexible JSON to compact binary formats, zero‑copy techniques, and selective compression to keep costs under control at tens of millions of requests per second.

FlatBuffersJSONProtobuf
0 likes · 15 min read
Why High‑QPS Systems Move from JSON to Efficient Binary Serialization
Random Bulletin
Random Bulletin
Jul 9, 2026 · Backend Development

From Short Connections to Connection Pools: Managing Millions of QPS

A midnight alert revealed latency jumping from 20 ms to 200 ms despite healthy CPU and memory, leading to an investigation that uncovered TCP connection handling—handshake delays, TIME_WAIT buildup, and CPU overhead—as the hidden bottleneck, and explains why short connections must evolve into pooled, globally managed connections at tens of millions of QPS.

TCPconnection managementconnection pool
0 likes · 18 min read
From Short Connections to Connection Pools: Managing Millions of QPS
Random Bulletin
Random Bulletin
Jul 9, 2026 · Backend Development

API Design for Ten‑Million QPS: From Generic to Specialized, BFF vs GraphQL

When a single generic endpoint that returns over thirty fields is stressed at ten‑million queries per second, it becomes the weakest link, exposing costs of over‑fetching, downstream amplification, cache inefficiency and coupling; the article dissects these issues and shows how specialized APIs, BFF layers, and GraphQL can address them, while weighing their trade‑offs.

API DesignBFFGraphQL
0 likes · 18 min read
API Design for Ten‑Million QPS: From Generic to Specialized, BFF vs GraphQL
Random Bulletin
Random Bulletin
Jul 7, 2026 · Backend Development

Managing API Versions at Ten‑Million QPS: From Single Version to Coexisting Multi‑Versions

A tiny field‑addition and rename in an order service caused a cascade failure in a three‑year‑old reconciliation job, illustrating how any incompatible change in a system handling tens of millions of QPS can trigger an avalanche, and the article explains why version management, compatible design, multi‑version coexistence, and disciplined governance are essential to keep such large‑scale services reliable.

API VersioningService Governancecompatible design
0 likes · 18 min read
Managing API Versions at Ten‑Million QPS: From Single Version to Coexisting Multi‑Versions
Random Bulletin
Random Bulletin
Jul 5, 2026 · Backend Development

From Ad‑hoc to Strict Layering: Managing Service Dependencies at Ten‑Million QPS

An unexpected restart of an edge user‑tag service caused a core transaction chain to fail, exposing how reverse and circular dependencies can turn a clean layered architecture into a tangled web; the article explains the hidden costs of such chaos and outlines two hard rules and a step‑by‑step path to enforce strict, acyclic layering for systems handling tens of millions of QPS.

circular dependencydependency governancehigh QPS
0 likes · 15 min read
From Ad‑hoc to Strict Layering: Managing Service Dependencies at Ten‑Million QPS
Random Bulletin
Random Bulletin
Jul 1, 2026 · Operations

Evolving Message Expiration for 10 Million QPS: From No TTL to Smart Policies

When a high‑traffic system processes billions of messages per day, stale “zombie” messages can corrupt business state; this article walks through the evolution from never‑expiring queues to uniform TTL, then per‑topic and per‑message TTL, and finally to smart, context‑aware expiration, detailing the engineering trade‑offs, implementation patterns, and operational checklist needed for reliable 10 M‑QPS message pipelines.

KafkaRocketMQdistributed systems
0 likes · 19 min read
Evolving Message Expiration for 10 Million QPS: From No TTL to Smart Policies
Random Bulletin
Random Bulletin
Jun 29, 2026 · Operations

Backlog Digestion at Ten‑Million QPS: From Adding Machines to Intelligent Scheduling

The article dissects how to handle massive message backlog in ten‑million‑QPS systems, explaining why simply adding consumer machines fails, and walks through six evolutionary stages—from manual scaling and auto‑scaling to traffic tiering, consumer‑side optimizations, dynamic strategies, and AI‑driven intelligent scheduling—while highlighting design trade‑offs, pitfalls, and practical tooling.

auto-scalingbacklog digestionhigh QPS
0 likes · 22 min read
Backlog Digestion at Ten‑Million QPS: From Adding Machines to Intelligent Scheduling
Random Bulletin
Random Bulletin
Jun 28, 2026 · Operations

Backlog Monitoring at 10 Million QPS: From Manual Checks to SLO‑Driven Automation

At 10 million QPS, message backlog can silently grow for hours, turning minutes of lag into hundreds of millions of undelivered messages; this article walks through a six‑stage evolution—from manual command‑line checks to SLO‑driven predictive monitoring—detailing metrics, pitfalls, and tool choices for a robust, multi‑dimensional alert system.

KafkaSLOalerting
0 likes · 21 min read
Backlog Monitoring at 10 Million QPS: From Manual Checks to SLO‑Driven Automation
Random Bulletin
Random Bulletin
Jun 27, 2026 · Operations

Cross‑Data‑Center Replication at Ten‑Million QPS: From Async to Semi‑Sync and How to Choose

The article examines why cross‑datacenter replication must evolve from simple asynchronous copying to semi‑synchronous and layered strategies at the ten‑million‑QPS scale, detailing latency, RPO, bandwidth costs, failover complexities, and practical selection guidelines for each business tier.

Failoverasynchronous replicationbandwidth optimization
0 likes · 17 min read
Cross‑Data‑Center Replication at Ten‑Million QPS: From Async to Semi‑Sync and How to Choose
Random Bulletin
Random Bulletin
Jun 26, 2026 · Operations

Message Duplication Is Inevitable: Building a Multi‑Layer Idempotency Middleware for Ten‑Million QPS

Message queues guarantee at‑least‑once delivery, making duplicate messages a normal feature; the article examines a real coupon‑distribution incident, critiques business‑level idempotency approaches, and outlines a layered platform‑wide middleware design—including unique keys, state machines, storage choices, and TTL strategies—to achieve reliable processing at ten‑million QPS scale.

distributed systemshigh QPSidempotency
0 likes · 19 min read
Message Duplication Is Inevitable: Building a Multi‑Layer Idempotency Middleware for Ten‑Million QPS
Cloud Architecture
Cloud Architecture
Jun 24, 2026 · Backend Development

Four Levels of Concurrency Control: Optimistic Locks to Message Queues for Million‑QPS

High‑concurrency systems must go beyond simple locking; the article breaks down four progressive strategies—optimistic locking, database pessimistic locking, Redis distributed locks, and finally message‑queue serialization—explaining their trade‑offs, implementation details, pitfalls, and how to combine them into a layered architecture that sustains million‑QPS workloads with consistency, throughput, and recoverability.

Optimistic LockPessimistic LockRedis Lock
0 likes · 32 min read
Four Levels of Concurrency Control: Optimistic Locks to Message Queues for Million‑QPS
Random Bulletin
Random Bulletin
Jun 24, 2026 · Operations

Rack Awareness: From Zero to High‑QPS – Boost Availability, Cut Cross‑AZ Traffic

A real‑world rack‑power outage showed that three‑replica Kafka clusters can lose all replicas when brokers share a failure domain, prompting a deep dive into rack awareness—how fault‑domain tags are injected, replica‑placement and leader‑distribution algorithms, consumer‑proximity reads, bandwidth costs, failure scenarios, and the stepwise evolution from hundred‑thousand to ten‑million QPS deployments.

KafkaPulsarRack Awareness
0 likes · 26 min read
Rack Awareness: From Zero to High‑QPS – Boost Availability, Cut Cross‑AZ Traffic
Random Bulletin
Random Bulletin
Jun 18, 2026 · Operations

Zero Message Loss at 10M QPS: From Best‑Effort to Proven Guarantees

The article dissects why message loss is an end‑to‑end engineering challenge across production, broker, and consumer stages, presents real‑world failure cases at ten‑million QPS, and outlines concrete strategies—ack loops, outbox tables, broker replication settings, consumer handling rules, observability, reconciliation, and SLA‑driven compensation—to evolve from best‑effort to provable, recoverable reliability.

Outbox PatternSLAhigh QPS
0 likes · 17 min read
Zero Message Loss at 10M QPS: From Best‑Effort to Proven Guarantees
Random Bulletin
Random Bulletin
Jun 16, 2026 · Backend Development

Consumer Rate Limiting: From Zero to Full Control in Million‑QPS Architectures

The article explains why consumer‑side rate limiting is essential in million‑QPS systems, detailing how unchecked consumers can overwhelm downstream services, and presents practical strategies—including pause/resume, token‑bucket algorithms, adaptive thresholds, and global coordination—to safely throttle consumption without dropping messages.

KafkaRate Limitingconsumer optimization
0 likes · 16 min read
Consumer Rate Limiting: From Zero to Full Control in Million‑QPS Architectures
Random Bulletin
Random Bulletin
Jun 10, 2026 · Backend Development

Message Compression at Scale: From No Compression to Selective Strategies

A midnight bandwidth alarm triggers a deep dive into message compression, revealing how algorithm choice, compression placement, batch handling, and selective policies evolve from simple no‑compression setups to multi‑layered strategies that balance CPU, bandwidth, disk costs, and latency in million‑QPS systems.

KafkaZstdbatch compression
0 likes · 20 min read
Message Compression at Scale: From No Compression to Selective Strategies
Random Bulletin
Random Bulletin
Jun 3, 2026 · Operations

Designing Partitions for Millions of QPS: From Default Settings to Precise Capacity Planning

The article explains why partition count is not a simple integer but the solution of a multi‑constraint capacity‑planning equation, walks through three design stages—from naïve defaults, through CPU‑core‑based heuristics, to a precise engineering model that balances throughput, ordering, scaling, fault domains and storage costs for million‑plus QPS workloads.

Capacity PlanningKafkaPartition Design
0 likes · 32 min read
Designing Partitions for Millions of QPS: From Default Settings to Precise Capacity Planning
Su San Talks Tech
Su San Talks Tech
Jun 12, 2025 · Information Security

Defending Against Million‑QPS Attacks: Rate Limiting, Fingerprinting, and Dynamic Rules

This article explains why a million‑QPS flood can cripple systems, outlines attackers' tactics, and presents a three‑layer defense strategy—including gateway rate limiting with Nginx + Lua, distributed circuit breaking via Sentinel, device fingerprinting, behavior analysis, and a dynamic rule engine—to protect high‑traffic services.

DDoSRate Limitingbehavior analysis
0 likes · 14 min read
Defending Against Million‑QPS Attacks: Rate Limiting, Fingerprinting, and Dynamic Rules
Alimama Tech
Alimama Tech
May 12, 2025 · Artificial Intelligence

Universal Recommendation Model (URM): A General Large‑Model Recall System for Advertising

The article presents the Universal Recommendation Model (URM), a large‑language‑model‑based recall framework that integrates world knowledge and e‑commerce expertise through knowledge injection and prompt‑driven alignment, achieving significant offline recall gains and a 3.1% increase in ad consumption while meeting high‑QPS, low‑latency production constraints.

AdvertisingMultimodalRecall
0 likes · 17 min read
Universal Recommendation Model (URM): A General Large‑Model Recall System for Advertising
Xiaohongshu Tech REDtech
Xiaohongshu Tech REDtech
Nov 7, 2024 · Artificial Intelligence

RTAMS-GANNS: A Real-Time Adaptive Multi-Stream GPU System for Online Approximate Nearest Neighbor Search

RTAMS‑GANNS, the award‑winning real‑time adaptive multi‑stream GPU system for online approximate nearest neighbor search, eliminates costly memory allocations and serial execution by using a dynamic memory‑block insertion algorithm and separate CUDA streams, cutting latency by 40‑80% and reliably serving over 100 million daily users in production.

Approximate Nearest NeighborGPUVector Insertion
0 likes · 19 min read
RTAMS-GANNS: A Real-Time Adaptive Multi-Stream GPU System for Online Approximate Nearest Neighbor Search
Architect's Guide
Architect's Guide
Aug 24, 2022 · Backend Development

Optimizing Long‑Connection Services with Netty: From Millions of Connections to High QPS

This article summarizes the challenges and optimization techniques for building a high‑performance long‑connection service with Netty, covering non‑blocking I/O, Linux kernel tuning, client‑side testing, VM‑based scaling, data‑structure tweaks, CPU and GC bottlenecks, and the final results of achieving hundreds of thousands of connections and tens of thousands of QPS on a single server.

GC TuningJava NIOLinux Tuning
0 likes · 14 min read
Optimizing Long‑Connection Services with Netty: From Millions of Connections to High QPS
Xiao Lou's Tech Notes
Xiao Lou's Tech Notes
May 25, 2022 · Backend Development

Why Caching Timestamps Can Slash CPU Usage in High‑QPS Java and Go Services

This article explores how naive timestamp retrieval can become a CPU bottleneck under high concurrency, demonstrates cache‑based optimizations used in Alibaba's Cobar and Sentinel projects, presents benchmark results, and proposes an adaptive algorithm to enable or disable caching based on real‑time QPS.

Javahigh QPStimestamp caching
0 likes · 11 min read
Why Caching Timestamps Can Slash CPU Usage in High‑QPS Java and Go Services
Xianyu Technology
Xianyu Technology
Apr 13, 2022 · Big Data

Real-time Multi-system Data Aggregation for Fan Tag System

The Xianyu fan‑tag system solves the challenge of displaying full‑history purchase counts with real‑time updates and low‑latency, high‑throughput queries by daily exporting multi‑system data to a LevelDB‑based KV store, converting schemas, and applying real‑time compensation from transaction and follow‑change messages, merging offline and live data to produce sorted fan lists at ~10 k QPS.

KV storagedata aggregationhigh QPS
0 likes · 6 min read
Real-time Multi-system Data Aggregation for Fan Tag System
ByteFE
ByteFE
Apr 11, 2022 · Backend Development

ByteDance Wallet Asset Middle Platform Design for 2022 Spring Festival High‑Traffic Reward Distribution

This article details ByteDance's wallet asset middle platform designed for the 2022 Spring Festival, covering eight‑app reward interoperability, high‑QPS challenges, token‑based asynchronous入账, budget control, stability measures, and fund‑safety guarantees, and includes practical solutions for hot‑key handling, budget throttling, and multi‑stage activity isolation.

ByteDanceFund SafetyToken Scheme
0 likes · 22 min read
ByteDance Wallet Asset Middle Platform Design for 2022 Spring Festival High‑Traffic Reward Distribution
FunTester
FunTester
Apr 19, 2021 · Operations

How to Add a Soft‑Start Mechanism for High‑QPS Performance Testing in Java

This article explains the concept of soft‑start in performance testing, presents Java implementations for both fixed‑thread and fixed‑QPS models, discusses error‑impact considerations, and provides practical code snippets to gradually ramp up load and improve measurement accuracy for high‑throughput services.

ConcurrencyJavaLoad Testing
0 likes · 8 min read
How to Add a Soft‑Start Mechanism for High‑QPS Performance Testing in Java
Ctrip Technology
Ctrip Technology
Feb 25, 2021 · Backend Development

Design and Implementation of a Cache Access Component and Update Platform for High‑QPS Scenarios

This article describes a backend architecture for a high‑traffic e‑commerce project, detailing a cache access component and a cache update platform that use asynchronous messaging, hotspot‑key handling, versioned cache entries, and Redis to achieve low latency, high QPS support and strong data consistency.

BackendCachingdistributed-systems
0 likes · 18 min read
Design and Implementation of a Cache Access Component and Update Platform for High‑QPS Scenarios
Meituan Technology Team
Meituan Technology Team
Dec 20, 2018 · Backend Development

Design and Performance Optimization of LruCache in Meituan DSP System

Meituan’s DSP system boosted high‑QPS ad serving performance by layering an LRU cache in front of Redis, then adding time‑based eviction, sharding the cache into HashLruCache instances to cut lock contention, and employing a zero‑copy, reference‑counted design, ultimately cutting average latency to about 20 % of the original and similarly reducing 99.9th‑percentile delays.

HashLruCacheLRUCacheMeituan DSP
0 likes · 15 min read
Design and Performance Optimization of LruCache in Meituan DSP System