Tagged articles

monitoring

2386 articles · Page 1 of 24
Linux Tech Enthusiast
Linux Tech Enthusiast
Sep 24, 2026 · Operations

18 Production-Ready Shell Scripts for Daily Sysadmin Tasks

This article presents 18 practical shell scripts for common system administration tasks including file consistency checks, log rotation, network traffic monitoring, FTP downloads, user input validation, Nginx 502 error auto-recovery, port scanning, batch file renaming, IP blocking, and IP address validation.

Automationbashfile operations
0 likes · 36 min read
18 Production-Ready Shell Scripts for Daily Sysadmin Tasks
Efficient Ops
Efficient Ops
Sep 23, 2026 · Operations

Uptime Kuma Review: 77.7K Stars, Lightweight Self-Hosted Monitoring Setup

This article reviews Uptime Kuma, a lightweight open-source self-hosted monitoring tool with 77.7K GitHub stars, detailing its multi-protocol monitoring, 90+ notification channels, 20-second check intervals, and providing step-by-step installation guides for Docker Compose and manual Node.js deployments with PM2 process management.

DockerNode.jsPM2
0 likes · 5 min read
Uptime Kuma Review: 77.7K Stars, Lightweight Self-Hosted Monitoring Setup
LuTiao Programming
LuTiao Programming
Sep 21, 2026 · Backend Development

Distributed Rate Limiting with Redis + Lua: Surviving API Floods in Spring Boot

After an external system hammered a Spring Boot search endpoint causing database connection exhaustion, the author builds a distributed fixed-window rate limiter using Redis and Lua for atomicity, wraps it with an annotation-driven AOP aspect supporting user, IP, API key, and global dimensions, returns proper HTTP 429 with Retry-After, discusses fixed-window limitations versus token bucket and sliding window, and covers fail-open/fail-closed strategies for Redis outages plus monitoring metrics.

AOPFail-OpenFixed Window
0 likes · 17 min read
Distributed Rate Limiting with Redis + Lua: Surviving API Floods in Spring Boot
Random Bulletin
Random Bulletin
Sep 19, 2026 · Operations

Second-Level Fault Detection at 10M QPS: Layered Signals & Safe Automation

This article explains how to reduce fault detection latency from minutes to seconds in 10M QPS systems by implementing layered signals, combined evidence detection, distributed judgment, event normalization, and safe automation guardrails, rather than simply increasing sampling frequency.

AlertingSLOdistributed systems
0 likes · 36 min read
Second-Level Fault Detection at 10M QPS: Layered Signals & Safe Automation
Golang Shines
Golang Shines
Sep 19, 2026 · Operations

SEAL Methodology for Production Troubleshooting: Veteran Ops Toolbox & Case Studies

A 10-year operations veteran shares the SEAL troubleshooting framework (Symptom, Environment, Analysis, Location), a curated toolbox (Prometheus, ELK, perf, tcpdump), real-world case studies (Redis avalanche, MySQL slow queries), incident grading, automation scripts, performance tuning, container/Kubernetes diagnostics, monitoring models, chaos engineering, and AIOps trends.

AIOpsAutomationPerformance Optimization
0 likes · 20 min read
SEAL Methodology for Production Troubleshooting: Veteran Ops Toolbox & Case Studies
liandk
liandk
Sep 18, 2026 · Backend Development

Hidden GC Pauses Causing Random API Timeouts: JVM STW Troubleshooting Guide

This article reveals how long GC stop-the-world pauses cause random API timeouts without errors, detailing two fault scenarios, a four-step GC log analysis SOP with specific JVM flags, four root causes including young generation sizing and memory fragmentation, and optimization strategies from emergency restarts to G1/ZGC upgrades with monitoring alerts.

AlertingFull GCG1
0 likes · 15 min read
Hidden GC Pauses Causing Random API Timeouts: JVM STW Troubleshooting Guide
Raymond Ops
Raymond Ops
Sep 14, 2026 · Operations

RocketMQ Production Operations: Cluster Setup, Retry Mechanisms & Dead Letter Queue Solutions

This comprehensive guide covers RocketMQ production operations including cluster deployment with NameServer and Broker configurations, message retry mechanisms with backoff strategies, dead letter queue handling and reprocessing, monitoring with Prometheus alerts, and troubleshooting procedures for common issues like message accumulation, disk full, and broker failures.

Cluster DeploymentDead Letter QueueMessage Queue
0 likes · 67 min read
RocketMQ Production Operations: Cluster Setup, Retry Mechanisms & Dead Letter Queue Solutions
Raymond Ops
Raymond Ops
Sep 5, 2026 · Operations

Linux Time Synchronization Mastery: Chrony Best Practices for Production Systems

Comprehensive guide covering Linux time concepts, NTP protocol, chrony vs ntpd, clock source selection, leap second handling, configuration templates for cloud, containers, Kubernetes, and isolated networks, plus verification, monitoring, troubleshooting, compliance automation, and rollback strategies.

KubernetesNTPTroubleshooting
0 likes · 54 min read
Linux Time Synchronization Mastery: Chrony Best Practices for Production Systems
MaGe Linux Operations
MaGe Linux Operations
Sep 5, 2026 · Operations

Troubleshooting High Redis Client Connections: A Step-by-Step Guide to Identification and Optimization

This comprehensive guide details a systematic approach to diagnosing and resolving high Redis client connection counts, covering connection models, configuration tuning, CLI analysis commands, application-level connection pool fixes, temporary mitigation tactics, and long-term monitoring best practices with real-world case examples.

CLIENT LISTJedisRedis
0 likes · 35 min read
Troubleshooting High Redis Client Connections: A Step-by-Step Guide to Identification and Optimization
Random Bulletin
Random Bulletin
Sep 2, 2026 · Operations

Automating Root‑Cause Analysis for Million‑QPS Systems: From Manual to AI‑Assisted

When a transaction‑success rate dropped at 02:13 AM and 186 alerts flooded the on‑call channel, engineers struggled to piece together fragmented evidence, highlighting why manual root‑cause analysis is slow at scale and how an evidence‑driven automated pipeline can narrow investigation space, rank candidates with confidence, and keep humans in the loop for safe remediation.

Automationincident responselarge-scale systems
0 likes · 26 min read
Automating Root‑Cause Analysis for Million‑QPS Systems: From Manual to AI‑Assisted
Random Bulletin
Random Bulletin
Aug 30, 2026 · Operations

Alert Tiering: From a Single Level to a P0‑P3 Multi‑Level Response System

The article examines why a single‑level alerting approach fails at massive scale, outlines the five practical limitations it creates, and presents a step‑by‑step framework—classification matrix, routed channels, SLA timers, escalation policies, and on‑call discipline—to allocate limited human attention to the most business‑critical incidents.

AlertingSLAalert tiering
0 likes · 18 min read
Alert Tiering: From a Single Level to a P0‑P3 Multi‑Level Response System
Random Bulletin
Random Bulletin
Aug 29, 2026 · Operations

Alert Convergence at Scale: From Simple Deduplication to AI‑Driven Clustering

A 40‑second DB jitter triggered over 3,000 alerts, but by applying a five‑layer alert‑convergence strategy—deduplication, grouping, inhibition & silencing, dependency‑based aggregation, and AI‑powered clustering—teams can reduce noise by up to 90 %, turning a storm of notifications into a single actionable signal.

AIOpsAlertingdeduplication
0 likes · 19 min read
Alert Convergence at Scale: From Simple Deduplication to AI‑Driven Clustering
Java Companion
Java Companion
Aug 28, 2026 · Operations

Beszel: A Sub‑10 MB Open‑Source Monitoring Panel That Can Replace Prometheus

Beszel is a lightweight, open‑source server‑monitoring dashboard written in Go that runs a Hub‑Agent architecture, offers per‑machine history charts, Docker container stats, alerting, multi‑user access and backup, and can be deployed in about ten minutes with Docker Compose, making it ideal for small‑scale environments where Prometheus feels heavyweight.

AlertingBeszelDocker Compose
0 likes · 9 min read
Beszel: A Sub‑10 MB Open‑Source Monitoring Panel That Can Replace Prometheus
Raymond Ops
Raymond Ops
Aug 25, 2026 · Operations

30 Essential Linux Ops Commands Every Engineer Should Know

This guide presents 30 of the most frequently used Linux commands for operations engineers, organized into seven categories and illustrated with real‑world scenarios, parameter examples, safety warnings, and step‑by‑step usage tips to help you manage files, monitor systems, handle processes, diagnose networks, control users, compress data, and manage services.

LinuxNetworkingShell
0 likes · 50 min read
30 Essential Linux Ops Commands Every Engineer Should Know
MaGe Linux Operations
MaGe Linux Operations
Aug 22, 2026 · Operations

Essential New Metrics for Monitoring MCP and Tool Calls in API Gateways

The article analyzes how the emergence of MCP, function calling, and agent toolchains transforms API gateway traffic, identifies blind spots in traditional monitoring, and proposes a three‑layer metric system—including request, inference, and tool‑call dimensions—along with concrete Prometheus metrics, alert rules, and implementation guidelines for reliable observability.

API GatewayMCPPrometheus
0 likes · 33 min read
Essential New Metrics for Monitoring MCP and Tool Calls in API Gateways
Coder Trainee
Coder Trainee
Aug 21, 2026 · Operations

When Logs Fill the Disk: How I Cleared Three Days of Log Files

A production server hit 100% disk usage, prompting the author to use df, du and find to locate oversized log files, uncover missing rotation, excessive debug logging and a looping exception, then perform urgent cleanup and implement logrotate, log level adjustments, and Prometheus alerts to prevent recurrence.

LinuxLog ManagementTroubleshooting
0 likes · 7 min read
When Logs Fill the Disk: How I Cleared Three Days of Log Files
Ray's Galactic Tech
Ray's Galactic Tech
Aug 20, 2026 · Backend Development

Payment Reconciliation: Detecting Lost Orders, Over‑payments and Auto‑Repairing Them

The article explains why successful payment does not guarantee correct accounting, defines discrepancy types such as lost orders, over‑payments and shortfalls, and presents a production‑grade reconciliation architecture with T+1 bill ingestion, shard‑based diff detection, state‑machine driven error handling, automatic repair workflows, and comprehensive monitoring.

auto-repairlost ordersmonitoring
0 likes · 45 min read
Payment Reconciliation: Detecting Lost Orders, Over‑payments and Auto‑Repairing Them
Efficient Ops
Efficient Ops
Aug 19, 2026 · Operations

8 Must-Have MCP Ops Components That Dramatically Boost Efficiency

The article introduces eight essential MCP components—Grafana, Jenkins, K8s, Playwright, GitHub, Zabbix, Prometheus, and Alibaba Cloud—detailing how each enhances monitoring, automation, resource management, and performance optimization to cut fault‑resolution time, lower manual effort, and improve system stability.

AutomationGrafanaKubernetes
0 likes · 7 min read
8 Must-Have MCP Ops Components That Dramatically Boost Efficiency
Ops Community
Ops Community
Aug 19, 2026 · Operations

Why Prometheus Metrics Have High Cardinality and How to Fix It

The article explains why Prometheus metric cardinality explodes, how it impacts memory, storage and query performance, and provides a step‑by‑step troubleshooting guide with concrete examples, code snippets, mitigation strategies, validation methods, and best‑practice recommendations for SREs.

AlertingPrometheusTSDB
0 likes · 24 min read
Why Prometheus Metrics Have High Cardinality and How to Fix It
Random Bulletin
Random Bulletin
Aug 18, 2026 · Operations

From Single‑Node to Distributed Monitoring Storage: Scaling TSDB for Million‑QPS

The article explains how a single‑machine Prometheus TSDB initially works well but eventually hits five scalability walls—capacity, single‑point failure, short retention, lack of global view, and throughput limits—and then details the step‑by‑step evolution to remote_write with object‑storage‑backed Thanos and finally to native distributed TSDBs such as Cortex, Mimir, and VictoriaMetrics, including their trade‑offs, costs, and practical selection guidance.

CortexPrometheusTSDB
0 likes · 21 min read
From Single‑Node to Distributed Monitoring Storage: Scaling TSDB for Million‑QPS
Linux Tech Enthusiast
Linux Tech Enthusiast
Aug 18, 2026 · Operations

Production Incident Troubleshooting Framework and Toolbox: Veteran Ops Engineer’s Real‑World Tips

A seasoned operations veteran shares a step‑by‑step incident‑response workflow, the SEAL troubleshooting methodology, essential monitoring and debugging tools, real‑world case studies, automated scripts, and best‑practice guidelines to help engineers quickly diagnose and resolve production outages.

AutomationLinuxincident management
0 likes · 16 min read
Production Incident Troubleshooting Framework and Toolbox: Veteran Ops Engineer’s Real‑World Tips
Random Bulletin
Random Bulletin
Aug 16, 2026 · Operations

Real‑Time Alerting at Million‑QPS: From Static Thresholds to Intelligent Detection

At massive scales of millions of QPS and hundreds of metrics, static alert thresholds become noisy and hard to maintain; the article walks through a stepwise evolution—adding sustained‑duration and multi‑condition rules, adopting dynamic baselines, leveraging anomaly detection, and applying correlation, RCA and SLO burn‑rate techniques—to transform alerts from simple threshold breaches into precise, user‑experience‑focused notifications while combating alert fatigue.

AIOpsAlertingSLO
0 likes · 18 min read
Real‑Time Alerting at Million‑QPS: From Static Thresholds to Intelligent Detection
Linyb Geek Road
Linyb Geek Road
Aug 15, 2026 · Operations

Key Metrics Every Ops Engineer Should Monitor

This article enumerates essential operational metrics—such as CPU, memory, disk and network I/O, response time, throughput, error rates, availability, MTBF/MTTR, security logs, and capacity‑planning indicators—explaining their meanings and recommended target values to help engineers comprehensively monitor system performance, stability, and efficiency.

availabilitycapacity planninglogging
0 likes · 10 min read
Key Metrics Every Ops Engineer Should Monitor
Random Bulletin
Random Bulletin
Aug 14, 2026 · Operations

Why Metric Counts Explode from Thousands to Millions in High‑QPS Systems

A single high‑cardinality label can cause a monitoring system to crash as metric series jump from hundreds of thousands to millions, overwhelming storage, queries, collection, cost, and signal‑to‑noise; the article explains the root cause, impact dimensions, and practical mitigation strategies such as cardinality control, distributed TSDBs, downsampling, pre‑aggregation, and proper division of metrics, traces, and logs.

PrometheusTSDBhigh cardinality
0 likes · 17 min read
Why Metric Counts Explode from Thousands to Millions in High‑QPS Systems
Golang Shines
Golang Shines
Aug 13, 2026 · Backend Development

Cut page load from 5 s to 500 ms: 12 Nginx tuning parameters

The article explains that reducing page load from five seconds to 500 ms requires more than just tweaking a few Nginx directives, outlines a systematic workflow—baseline measurement, hypothesis, single‑parameter gray rollout, verification and rollback—and details twelve specific Nginx settings that can eliminate proven Web‑layer bottlenecks such as file‑descriptor limits, connection models, static‑file paths, compression and request buffering.

LinuxNginxmonitoring
0 likes · 26 min read
Cut page load from 5 s to 500 ms: 12 Nginx tuning parameters
Java Tech Workshop
Java Tech Workshop
Aug 13, 2026 · Backend Development

Is Your SpringBoot @Scheduled Task Reliable? A Full‑Stack Breakdown

This article examines the hidden pitfalls of SpringBoot’s @Scheduled annotation—such as duplicate runs in clusters, single‑thread blocking, uncaught exceptions, and lack of monitoring—and provides a step‑by‑step guide to configuring custom thread pools, distributed locks, timeout handling, dynamic task management, and observability for production‑grade reliability.

Redis LockSpringBootThreadPool
0 likes · 19 min read
Is Your SpringBoot @Scheduled Task Reliable? A Full‑Stack Breakdown
Golang Shines
Golang Shines
Aug 12, 2026 · Operations

Boost Ops Efficiency: 10 Essential Linux Tools Every Engineer Should Use

This article presents a practical guide for system administrators and DevOps engineers, introducing ten high‑frequency Linux tools—htop, iotop, nethogs, ncdu, strace, lsof, tcpdump, netstat/ss, curl, and systemctl/journalctl—detailing their installation, core and advanced usage, real‑world case studies, and how to combine them to dramatically improve troubleshooting speed and overall operational efficiency.

Linuxhtopiotop
0 likes · 54 min read
Boost Ops Efficiency: 10 Essential Linux Tools Every Engineer Should Use
SpringMeng
SpringMeng
Aug 12, 2026 · Databases

RedisInsight: The Official High‑Performance GUI for Redis

This article introduces RedisInsight, the official visual management tool for Redis, outlines its key features, provides step‑by‑step installation on Linux and Kubernetes, and demonstrates basic usage for monitoring, querying, and memory analysis through the GUI.

GUIKubernetesLinux
0 likes · 7 min read
RedisInsight: The Official High‑Performance GUI for Redis
Coder Life Journal
Coder Life Journal
Aug 11, 2026 · Backend Development

Spring Cloud + Kafka: 6 Common Pitfalls and How to Avoid Them

The article walks through six real‑world pitfalls when integrating Spring Cloud with Kafka—message loss, duplicate processing, out‑of‑order events, massive consumer lag, serialization mismatches, and misuse of Kafka transactions—and provides concrete configuration tweaks, code examples, and operational safeguards to prevent each issue.

KafkaSpring Clouddistributed systems
0 likes · 9 min read
Spring Cloud + Kafka: 6 Common Pitfalls and How to Avoid Them
Raymond Ops
Raymond Ops
Aug 9, 2026 · Operations

How a Full Redis Connection Pool Triggered a Service Outage: Step‑by‑Step Investigation

An online education platform experienced a cascade failure when Redis reached its maxclients limit, causing authentication, session, and cache services to become unavailable; the article details the connection mechanism, root‑cause analysis, rapid mitigation steps, and long‑term best practices for preventing similar outages.

JedisRediscircuit breaker
0 likes · 18 min read
How a Full Redis Connection Pool Triggered a Service Outage: Step‑by‑Step Investigation
Alibaba Cloud Native
Alibaba Cloud Native
Aug 9, 2026 · Mobile Development

Reconstructing an AI App’s Waiting Experience with Flutter RUM Monitoring

This article explains how to use Alibaba Cloud's Flutter RUM SDK to correlate user actions, network requests, long‑tasks and errors, reconstructing the full “spinning page” scenario in AI applications, and shows how STAROps can pinpoint interface failures, client‑side blocks, and rendering bottlenecks with concrete integration steps and code examples.

FlutterSTAROpsmobile
0 likes · 20 min read
Reconstructing an AI App’s Waiting Experience with Flutter RUM Monitoring
Raymond Ops
Raymond Ops
Aug 8, 2026 · Operations

A Complete Walkthrough of Investigating High Server Load in Production

This article narrates a step‑by‑step investigation of a sudden CPU load spike on a 24‑core e‑commerce web server, revealing an I/O bottleneck caused by misconfigured log rotation and excessive debug logging, and outlines the diagnostic commands, root‑cause analysis, immediate remediation, and long‑term fixes.

I/O BottleneckLinuxLoad Average
0 likes · 19 min read
A Complete Walkthrough of Investigating High Server Load in Production
Ops Community
Ops Community
Aug 7, 2026 · Operations

Node Exporter Metrics Explained: CPU, Memory, Disk & Network Monitoring

This guide walks through a systematic investigation of Node Exporter metrics—starting with verifying the scrape pipeline, then analyzing CPU, memory, disk, and network data using PromQL queries, command‑line checks, and alert‑rule examples—to help operators pinpoint resource bottlenecks and configure reliable monitoring.

AlertingCPUPrometheus
0 likes · 25 min read
Node Exporter Metrics Explained: CPU, Memory, Disk & Network Monitoring
MaGe Linux Operations
MaGe Linux Operations
Aug 6, 2026 · Databases

How to Determine the Right Database Connection Pool Size: Practical Guidelines and Benchmarks

This article walks through a systematic approach to sizing PostgreSQL connection pools for Java applications using HikariCP and Spring Boot, covering capacity budgeting, workload‑driven calculations, monitoring metrics, slow‑SQL analysis, leak detection, Kubernetes deployment considerations, and safe rollout practices.

HikariCPKubernetesPostgreSQL
0 likes · 27 min read
How to Determine the Right Database Connection Pool Size: Practical Guidelines and Benchmarks
Xiaolin Talks Programming
Xiaolin Talks Programming
Aug 6, 2026 · Backend Development

Spring Boot + Quartz: Production-Grade Scheduling with Dynamic Control, Cluster HA & Misfire Handling

This guide details building a production-grade Spring Boot Quartz scheduler covering JDBC persistence, cluster coordination via database locks, misfire policies, dynamic job management, execution logging via JobListener, idempotency patterns, monitoring metrics, and troubleshooting common pitfalls like transaction boundaries and clock drift.

ClusterQuartzSpring Boot
0 likes · 20 min read
Spring Boot + Quartz: Production-Grade Scheduling with Dynamic Control, Cluster HA & Misfire Handling
Ops Community
Ops Community
Aug 5, 2026 · Operations

Linux Kernel Sysctl Tuning Checklist – Proven Steps to Improve Performance

This article debunks the myth that simply copying a sysctl.conf yields a 30% boost, and presents a rigorous engineering loop—baseline measurement, hypothesis formulation, gray‑scale changes, observation of side effects, and rollback—along with detailed scripts, metrics, and per‑parameter guidance for memory, network, file handles, and more.

Linuxkernel parametersmonitoring
0 likes · 37 min read
Linux Kernel Sysctl Tuning Checklist – Proven Steps to Improve Performance
Raymond Ops
Raymond Ops
Aug 4, 2026 · Information Security

A Miswritten iptables Rule That Almost Made Me Quit

The article recounts a real‑world iptables misconfiguration that cut off SSH access for 47 minutes, walks through the incident timeline, root‑cause analysis, and detailed remediation steps, and then expands into a comprehensive guide on iptables fundamentals, common pitfalls, best‑practice design, troubleshooting commands, automation, monitoring, and migration to nftables.

AutomationLinuxfirewall
0 likes · 71 min read
A Miswritten iptables Rule That Almost Made Me Quit
Linyb Geek Road
Linyb Geek Road
Aug 2, 2026 · Operations

How to Build a Systematic Enterprise Monitoring Architecture

This article outlines a comprehensive, step‑by‑step approach for constructing a systematic enterprise monitoring system, covering the four core technical modules (collection, data, operators, alerts), designing a layered metric framework, and establishing a health‑management lifecycle that includes proactive alert prevention, real‑time handling, and post‑incident review.

AlertingCMDBSRE
0 likes · 21 min read
How to Build a Systematic Enterprise Monitoring Architecture
Linyb Geek Road
Linyb Geek Road
Aug 2, 2026 · Operations

What Makes This Ops Expert’s Monitoring System Design So Effective?

The article explains how to build a comprehensive monitoring system using the USE method, outlines essential system and application metrics, and walks through the architecture and components of Prometheus, Grafana, full‑link tracing, and the ELK stack for effective operations monitoring.

ELKPrometheusUSE method
0 likes · 13 min read
What Makes This Ops Expert’s Monitoring System Design So Effective?
MaGe Linux Operations
MaGe Linux Operations
Jul 30, 2026 · Operations

Cut Page Load from 5 s to 500 ms: 12 Essential Nginx Performance Tweaks

The article explains how to reduce overall page latency from five seconds to half a second by systematically measuring, hypothesizing, and tuning twelve Nginx directives—such as worker processes, file‑descriptor limits, keep‑alive settings, and gzip—while backing up configurations, performing gray‑scale rollouts, and validating results with curl and log analysis.

LinuxNginxconfiguration
0 likes · 25 min read
Cut Page Load from 5 s to 500 ms: 12 Essential Nginx Performance Tweaks
Raymond Ops
Raymond Ops
Jul 28, 2026 · Databases

MySQL Disk Space Explodes: How Binary Logs Become the Hidden Culprit

MySQL servers can trigger alarming disk‑space warnings even when the data directory is small, because unchecked binary logs rapidly consume storage; this article explains the log’s purpose, why it grows, how to diagnose the issue, and provides step‑by‑step cleanup, configuration, replication, monitoring, and recovery best practices.

Binary LogMySQLReplication
0 likes · 26 min read
MySQL Disk Space Explodes: How Binary Logs Become the Hidden Culprit
samdeepthink
samdeepthink
Jul 28, 2026 · Backend Development

Why Ignoring Non‑Critical Alerts Is Safe When Core Services Are Monitored

The author explains that when core business modules are properly monitored with structured logs and dedicated alert groups, a flood of non‑critical alerts can be ignored, shares a lightweight Java StructuredLog utility, and outlines best practices for real‑time monitoring and log standardization.

DingTalkJavaStructured Logging
0 likes · 9 min read
Why Ignoring Non‑Critical Alerts Is Safe When Core Services Are Monitored
Raymond Ops
Raymond Ops
Jul 27, 2026 · Databases

What to Do First When MySQL Connections Are Maxed Out

This guide walks you through a complete emergency response, root‑cause analysis, and long‑term mitigation for MySQL connection‑limit exhaustion, covering Linux diagnostics, SQL commands, quick‑kill scripts, monitoring with Prometheus/Grafana, and best‑practice configuration of connection pools and max_connections.

MySQLTroubleshootingconnection limits
0 likes · 37 min read
What to Do First When MySQL Connections Are Maxed Out
Java Architect Handbook
Java Architect Handbook
Jul 27, 2026 · Backend Development

Druid Crashed in Production? Essential Optimizations for Spring Boot

The article explains why Druid connection pools can fail in production and provides a step‑by‑step guide to extreme optimization, covering environment setup, core pool parameter tuning, monitoring with StatFilter and web UI, security hardening, leak detection, dynamic adjustments, and common pitfalls.

Advanced OptimizationDruidSpring Boot
0 likes · 15 min read
Druid Crashed in Production? Essential Optimizations for Spring Boot
Ops Community
Ops Community
Jul 26, 2026 · Operations

When APM Agent Overheads Spike: In‑Depth Comparison of SkyWalking vs Pinpoint

The article provides a step‑by‑step methodology for establishing a performance baseline, verifying JVM and agent versions, collecting CPU, memory, GC, thread, and network metrics, configuring SkyWalking and Pinpoint agents, running controlled experiments, and using the results to choose the most suitable APM probe.

APMJavaPinpoint
0 likes · 15 min read
When APM Agent Overheads Spike: In‑Depth Comparison of SkyWalking vs Pinpoint
Linyb Geek Road
Linyb Geek Road
Jul 26, 2026 · Operations

From Alert Flood to Fault Insight: The Real Starting Point of AIOps

The article explains that successful AIOps begins not with sophisticated models but with turning a flood of fragmented alerts into a single, context‑rich incident view that tells operators how many failures occurred, which business services are impacted, and where they should start investigating.

AIOpsalert aggregationchange management
0 likes · 14 min read
From Alert Flood to Fault Insight: The Real Starting Point of AIOps
Linyb Geek Road
Linyb Geek Road
Jul 26, 2026 · Operations

Postmortem: How an Alert Flood Masked the Real Problem

A late‑night incident flooded the on‑call channel with dozens of red alerts, hiding the true root cause—a core service latency spike—until the team re‑ordered information, prioritized early signals, and applied a simple three‑tier alert classification to restore clarity and speed up resolution.

AIOpsalert managementincident response
0 likes · 11 min read
Postmortem: How an Alert Flood Masked the Real Problem
MaGe Linux Operations
MaGe Linux Operations
Jul 25, 2026 · Operations

How to Diagnose and Automate Cleanup for Server Disk‑Full Alerts

The guide explains why a full disk is more than just deleting a large file, walks through preserving evidence, checking capacity, inode and storage health, pinpointing the root cause on Linux hosts with systemd and Docker, and implementing safe automated cleanup with monitoring and rollback.

AutomationDockerLinux
0 likes · 22 min read
How to Diagnose and Automate Cleanup for Server Disk‑Full Alerts
Ops Development Stories
Ops Development Stories
Jul 25, 2026 · Cloud Native

Practical Guide to Pyrra: The Kubernetes‑Native SLO Monitoring Tool

This comprehensive guide explains how Pyrra extends Sloth by providing a full SLO platform for Kubernetes, covering its architecture, four SLI types, rule generation, Web UI features, alert configuration, deployment options, Grafana integration, advanced usage, common pitfalls, and a detailed comparison to help you choose the right tool for reliable service monitoring.

Cloud NativeKubernetesPrometheus
0 likes · 24 min read
Practical Guide to Pyrra: The Kubernetes‑Native SLO Monitoring Tool
MaGe Linux Operations
MaGe Linux Operations
Jul 23, 2026 · Operations

When Linux CPU spikes to 99%, these 6 commands saved me three times

If a Linux server shows 99% CPU usage, don’t restart immediately; instead use a systematic chain of six diagnostic commands—uptime, top, ps, pidstat, mpstat, and sar—combined with cgroup, thread, and log analysis to pinpoint the true cause, whether user‑space computation, I/O wait, soft‑interrupts, virtualization steal, or container limits, and then apply targeted remediation.

CPULinuxcgroup
0 likes · 24 min read
When Linux CPU spikes to 99%, these 6 commands saved me three times
Ops Community
Ops Community
Jul 21, 2026 · Operations

How to Extend Zabbix Without Writing Any Code

This article presents a code‑free Zabbix Agent deployment module that lets administrators batch‑install agents via the Zabbix web UI, explains its key features, typical use cases, step‑by‑step installation instructions, and showcases the resulting monitoring setup with screenshots.

Agent DeploymentAutomationCSV Import
0 likes · 7 min read
How to Extend Zabbix Without Writing Any Code
Wu Shixiong's Large Model Academy
Wu Shixiong's Large Model Academy
Jul 21, 2026 · Artificial Intelligence

How to Decompose a Production‑Ready RAG System for Interview Success

The article outlines a production‑ready RAG architecture by separating offline ingestion and online query pipelines, detailing nine ingestion steps, online request flow, data storage responsibilities, failure‑handling, monitoring, and acceptance criteria, all illustrated with concrete examples and traceable state machines.

RAGSystem Designfailure handling
0 likes · 29 min read
How to Decompose a Production‑Ready RAG System for Interview Success
CodeSmart Hoops
CodeSmart Hoops
Jul 19, 2026 · Interview Experience

Test Your Grafana Knowledge: 8 Interview Questions with Answers

This article provides a comprehensive Grafana guide covering core concepts, dashboard design principles, panel types, variable templating, alerting strategies, provisioning as code, high‑availability setup, and a multi‑region monitoring screen design, each illustrated with concrete examples and configuration snippets.

AlertingGrafanaPrometheus
0 likes · 26 min read
Test Your Grafana Knowledge: 8 Interview Questions with Answers
Raymond Ops
Raymond Ops
Jul 18, 2026 · Databases

MySQL Master‑Slave Replication: Core Architecture, GTID Setup, and Common Troubleshooting

This article provides a comprehensive, hands‑on guide to MySQL master‑slave replication, covering the underlying architecture, binlog formats, GTID and semi‑synchronous modes, detailed configuration steps, thread workflows, common failure scenarios with step‑by‑step diagnostics, and practical monitoring and failover scripts.

GTIDMySQLReplication
0 likes · 38 min read
MySQL Master‑Slave Replication: Core Architecture, GTID Setup, and Common Troubleshooting
TechVision Expert Circle
TechVision Expert Circle
Jul 17, 2026 · Artificial Intelligence

Building Trustworthy AI Systems: Core Dimensions and Practical Solutions

The article outlines a comprehensive engineering approach for trustworthy AI, detailing five measurable dimensions—safety, reliability, explainability, privacy, and fairness—along with architecture design, input/output safeguards, hallucination mitigation, monitoring metrics, human‑in‑the‑loop strategies, and real‑world trade‑off recommendations.

AI safetyLLM engineeringRAG
0 likes · 13 min read
Building Trustworthy AI Systems: Core Dimensions and Practical Solutions
Cloud Architecture
Cloud Architecture
Jul 17, 2026 · Backend Development

Elasticsearch Cluster: Inverted Index, Mechanics, Architecture for 100M+ Queries

This article provides a comprehensive, production‑grade guide to Elasticsearch clusters, covering the fundamentals of inverted indexes and Lucene segments, near‑real‑time write mechanics, shard routing, indexing pipelines, query execution flow, scaling strategies, and practical tips to avoid common pitfalls in high‑traffic search systems.

ElasticsearchJavaindexing
0 likes · 37 min read
Elasticsearch Cluster: Inverted Index, Mechanics, Architecture for 100M+ Queries
Raymond Ops
Raymond Ops
Jul 17, 2026 · Operations

Linux Disk Space Alerts? Locate the Problem in 3 Quick Steps

When a Linux disk space alert fires, this guide walks you through three rapid steps—identifying the affected partition, deep‑diving with df, du, ncdu and custom scripts, and cleaning up logs, caches, inodes, LVM, Docker, and quotas—to quickly pinpoint and resolve the root cause.

DockerLVMLinux
0 likes · 44 min read
Linux Disk Space Alerts? Locate the Problem in 3 Quick Steps
Su San Talks Tech
Su San Talks Tech
Jul 16, 2026 · Operations

How Skywalking’s New AI Features Turn Observability into Intelligent Diagnosis

Skywalking now integrates AI across three layers—Horizon UI AI Assistant for natural‑language queries, Virtual GenAI for transparent LLM call monitoring, and AI Pipeline for proactive, machine‑learning‑driven ops—providing step‑by‑step deployment guidance, real‑time chart generation, cost estimation, and use‑case recommendations.

AI PipelineApache SkywalkingHorizon UI
0 likes · 18 min read
How Skywalking’s New AI Features Turn Observability into Intelligent Diagnosis
Go Development Architecture Practice
Go Development Architecture Practice
Jul 14, 2026 · Operations

Embedded Monitoring Best Practice: Use go-commons for Built-in Service Health Reports

This article demonstrates how to quickly add lightweight, plug‑and‑play monitoring to a Go service using the open‑source go-commons library, showing installation, a minimal 50‑line example that exposes business QPS and system metrics via a single /metrics endpoint, and how to integrate it with Prometheus and Grafana.

GoPrometheusgo-commons
0 likes · 6 min read
Embedded Monitoring Best Practice: Use go-commons for Built-in Service Health Reports
samdeepthink
samdeepthink
Jul 14, 2026 · Backend Development

Thread‑Pool Outage Postmortem: Four Defense Layers to Prevent Data Loss

A July 13 incident revealed that sharing a single thread pool across order, refund, and status sync services caused queue saturation, task rejection, and data loss, prompting a four‑layer defense—pool isolation, CallerRunsPolicy with structured alerts, minute‑level DingTalk notifications, and a compensation tool—to ensure reliability and quick recovery.

AlertingJavaThread Pool
0 likes · 10 min read
Thread‑Pool Outage Postmortem: Four Defense Layers to Prevent Data Loss
Raymond Ops
Raymond Ops
Jul 13, 2026 · Operations

Scaling Prometheus to Thousands of Nodes with Thanos: Architecture, Storage, and HA Practices

The article analyzes the storage, query performance, high‑availability, and data‑loss challenges of running Prometheus on a 1,000‑node Kubernetes cluster and demonstrates how a Thanos‑based architecture—Sidecar, Query, Store Gateway, Compactor, Receiver, and object‑storage back‑ends—can be designed, tuned, and operated to achieve horizontal scalability, efficient down‑sampling, and reliable fault recovery.

KubernetesObject StoragePrometheus
0 likes · 35 min read
Scaling Prometheus to Thousands of Nodes with Thanos: Architecture, Storage, and HA Practices
Ops Community
Ops Community
Jul 13, 2026 · Operations

How to Diagnose a Suddenly Lagging Linux Server: Step‑by‑Step Ops Checklist

This guide walks you through a systematic, read‑only diagnostic workflow for a Linux server that becomes unresponsive, covering initial symptom clarification, data collection, CPU, memory, disk, network, application, container, and post‑mortem analysis, with concrete commands and evidence‑based decision points.

LinuxServerTroubleshooting
0 likes · 37 min read
How to Diagnose a Suddenly Lagging Linux Server: Step‑by‑Step Ops Checklist
Raymond Ops
Raymond Ops
Jul 12, 2026 · Operations

Essential Port Connectivity Troubleshooting: A Complete Step‑by‑Step Guide

This guide walks you through a systematic, seven‑layer approach to diagnosing port connectivity failures on Linux systems, covering service listening checks, local firewall rules, SELinux policies, network path analysis, cloud security groups, and application‑level protocols, with concrete commands, scripts, case studies, best‑practice recommendations, and monitoring tips.

LinuxSELinuxfirewall
0 likes · 40 min read
Essential Port Connectivity Troubleshooting: A Complete Step‑by‑Step Guide
Linyb Geek Road
Linyb Geek Road
Jul 12, 2026 · Operations

Designing a High‑Availability Architecture: Core Principles and Practices

This article outlines the essential principles for building a high‑availability system, covering cluster and distributed designs, fault‑tolerance, reliable hardware, disaster recovery, monitoring, security, capacity planning, and automated scaling to achieve optimal performance and resilience.

Automationcapacity planningcluster architecture
0 likes · 6 min read
Designing a High‑Availability Architecture: Core Principles and Practices
IT Learning Made Simple
IT Learning Made Simple
Jul 10, 2026 · Backend Development

Top 10 Architecture Design Mistakes and How to Avoid Them

This guide enumerates the ten most common architecture design mistakes—over‑design, ignoring business needs, single points of failure, premature optimization, chaotic tech stacks, tight coupling, missing monitoring, security oversights, and team capability gaps—explaining their symptoms, costly consequences, and concrete best‑practice remedies, plus checklists to keep your system robust and maintainable.

Backendarchitecturedesign pitfalls
0 likes · 11 min read
Top 10 Architecture Design Mistakes and How to Avoid Them
Black & White Path
Black & White Path
Jul 10, 2026 · Information Security

Five Repeating Mistakes Behind 50 API Leak Disasters

Analyzing over 50 major API breach cases, the article reveals that a single recurring vulnerability—Broken Object Level Authorization—combined with four systemic errors in trust management, secret handling, monitoring, and project‑based security thinking, repeatedly expose millions of users' sensitive data.

API SecurityBroken Object Level AuthorizationSecret Management
0 likes · 21 min read
Five Repeating Mistakes Behind 50 API Leak Disasters
CodeSmart Hoops
CodeSmart Hoops
Jul 10, 2026 · Cloud Native

Prometheus Deep Dive: The De Facto Standard for Cloud‑Native Monitoring

The article walks through a real‑world migration from Zabbix to Prometheus, explaining its pull‑based design, metric model, PromQL language, service discovery options, remote storage choices, alerting with Alertmanager, and a complete Spring Boot integration, while highlighting best‑practice recommendations and common pitfalls.

AlertmanagerCloud NativePromQL
0 likes · 34 min read
Prometheus Deep Dive: The De Facto Standard for Cloud‑Native Monitoring
Ops Community
Ops Community
Jul 8, 2026 · Operations

Quick Nginx Log Analysis Techniques to Spot Abnormal Requests and Attack Sources

This article provides a step‑by‑step guide on using Nginx's custom log_format together with command‑line tools such as awk, grep, sort and jq to identify slow requests, 5xx spikes, CC attacks, scanners and SQL‑injection attempts, and then mitigates them with limit_req, map, geo and iptables rules, while also covering log rotation, monitoring and risk‑aware deployment practices.

DevOpsNginxlog-analysis
0 likes · 37 min read
Quick Nginx Log Analysis Techniques to Spot Abnormal Requests and Attack Sources
Raymond Ops
Raymond Ops
Jul 8, 2026 · Operations

10 Essential System Commands Every Ops Engineer Should Master

This guide explains why core Linux commands are indispensable for ops engineers, categorizes them into monitoring, networking, disk analysis, text processing and service management, and provides detailed usage examples, troubleshooting scenarios, and decision‑tree guidance for effective system administration.

Linuxdisk managementmonitoring
0 likes · 47 min read
10 Essential System Commands Every Ops Engineer Should Master
Raymond Ops
Raymond Ops
Jul 7, 2026 · Operations

Practical Guide to Diagnosing and Resolving Linux Disk Space Exhaustion

This article provides a step‑by‑step, command‑driven methodology for identifying the five root causes of full disk space on Linux systems—block exhaustion, inode depletion, deleted‑but‑still‑held files, reserved space, and filesystem corruption—and offers concrete remediation techniques, automation scripts, and best‑practice recommendations.

InodeLVMLinux
0 likes · 55 min read
Practical Guide to Diagnosing and Resolving Linux Disk Space Exhaustion
Golang Shines
Golang Shines
Jul 7, 2026 · Operations

Mastering Linux Server Time Synchronization with NTP and Chrony: Best Practices

This guide explains why time synchronization is a critical yet often overlooked part of Linux operations, outlines common failure scenarios, and provides a step‑by‑step methodology for configuring, verifying, and troubleshooting NTP/chrony across physical servers, virtual machines, containers, and Kubernetes clusters.

KubernetesLinuxNTP
0 likes · 42 min read
Mastering Linux Server Time Synchronization with NTP and Chrony: Best Practices
Full-Stack DevOps & Kubernetes
Full-Stack DevOps & Kubernetes
Jul 6, 2026 · Cloud Native

Taming Massive Alert Noise: A Hands‑On Guide to AI‑Driven Dynamic Thresholds for Prometheus

This article presents a practical solution that uses Facebook Prophet time‑series AI to automatically calibrate dynamic alert thresholds in Prometheus, reducing over‑80% of false alarms in Kubernetes environments by learning business cycles and updating rules hourly without manual intervention.

AIOpsDynamic ThresholdFacebook Prophet
0 likes · 10 min read
Taming Massive Alert Noise: A Hands‑On Guide to AI‑Driven Dynamic Thresholds for Prometheus
Cloud Architecture
Cloud Architecture
Jul 5, 2026 · Databases

Production‑grade Elasticsearch: From Lucene Internals to Billion‑scale Search Architecture

This article explains why running Elasticsearch in production is far more complex than a simple API tutorial, covering Lucene fundamentals, mapping design, shard planning, write‑path architecture, query optimization, hot‑warm‑cold tiering, ILM policies, capacity planning, monitoring, incident handling, and a step‑by‑step evolution roadmap for building a reliable, scalable search system.

ElasticsearchIndex ModelingLucene
0 likes · 54 min read
Production‑grade Elasticsearch: From Lucene Internals to Billion‑scale Search Architecture
Raymond Ops
Raymond Ops
Jul 5, 2026 · Operations

Building a Basic Monitoring System from Zero: How to View CPU, Memory, Disk, and Network

This article walks you through setting up a complete monitoring stack with Prometheus, node_exporter, Grafana and Alertmanager, explains how to interpret the four core dimensions—CPU, memory, disk and network—using a structured troubleshooting workflow, and provides real‑world case studies, scripts and best‑practice recommendations.

AlertmanagerGrafanaLinux
0 likes · 35 min read
Building a Basic Monitoring System from Zero: How to View CPU, Memory, Disk, and Network
Golang Shines
Golang Shines
Jul 5, 2026 · Operations

Simple Foolproof Zabbix Deployment Guide

This step‑by‑step tutorial shows how to quickly install Zabbix using the VOF package for learning purposes, covering download, system and network configuration, and basic access testing, while warning that it is not suited for production environments.

LinuxVOFdeployment
0 likes · 3 min read
Simple Foolproof Zabbix Deployment Guide
Xiaolin Talks Programming
Xiaolin Talks Programming
Jul 5, 2026 · Backend Development

Spring Boot HTTP Client Tuning: From RestTemplate to HttpClient 5 for Production Connection Pools

This guide details production-grade HTTP client tuning in Spring Boot, covering RestTemplate pitfalls, HttpClient 5 configuration, connection pool sizing via Little's Law, timeout layering, HTTP/2 trade-offs, business isolation patterns, interceptor chains, graceful shutdown, and metrics-driven troubleshooting for connection leaks.

HTTP/2HttpClient 5RestTemplate
0 likes · 18 min read
Spring Boot HTTP Client Tuning: From RestTemplate to HttpClient 5 for Production Connection Pools
Raymond Ops
Raymond Ops
Jul 3, 2026 · Operations

10 Rookie Ops Mistakes You Must Avoid – A Complete Checklist

This guide walks ops newcomers through the ten most common pitfalls—from accidental rm‑rf deletions and mis‑configured firewalls to unsafe chmod usage—and provides concrete remediation steps, ready‑to‑run shell scripts, best‑practice checklists, and monitoring setups to keep production environments stable and secure.

DevOpsLinuxmonitoring
0 likes · 51 min read
10 Rookie Ops Mistakes You Must Avoid – A Complete Checklist
Ops Community
Ops Community
Jul 3, 2026 · Operations

10 Essential Shell Scripts to Halve Your Ops Workload

These ten practical Bash scripts automate common sysadmin tasks—disk space checks, log rotation, resource monitoring, backup validation, process guarding, port probing, and more—providing reusable, idempotent solutions with logging, alerting, dry‑run support, and cron integration to streamline operations.

AutomationCronShell
0 likes · 42 min read
10 Essential Shell Scripts to Halve Your Ops Workload
Raymond Ops
Raymond Ops
Jul 2, 2026 · Operations

How to Monitor Large Model Applications: A Beginner‑Friendly Metric System

This guide walks you through building a production‑grade monitoring solution for large language model inference services using a three‑layer metric hierarchy, Prometheus, Grafana, DCGM Exporter, and custom Python metrics, with step‑by‑step deployment, alerting policies, and real‑world troubleshooting examples.

AI InfrastructureGrafanaLarge Language Models
0 likes · 42 min read
How to Monitor Large Model Applications: A Beginner‑Friendly Metric System
Random Bulletin
Random Bulletin
Jun 30, 2026 · Operations

Preventing Message Queue Backlog at Ten‑Million QPS: From Reactive to Proactive

The article explains how a ten‑million‑QPS messaging system can shift from reactive fire‑fighting to proactive backlog prevention by using long‑, mid‑, and short‑term defense layers, capacity planning, predictive scaling, backpressure, and regular chaos engineering to eliminate user‑visible latency.

Message Queuebacklog preventionbackpressure
0 likes · 19 min read
Preventing Message Queue Backlog at Ten‑Million QPS: From Reactive to Proactive
Java Baker
Java Baker
Jun 29, 2026 · Backend Development

How to Diagnose Uneven CPU Usage in Java Services Using Kafka

This article walks through the symptoms, root cause analysis, and step‑by‑step solutions for uneven CPU usage across Java service instances, highlighting how mismatched Kafka partition counts and thread or GC issues can lead to load imbalance and how to resolve them.

CPUJavaKafka
0 likes · 8 min read
How to Diagnose Uneven CPU Usage in Java Services Using Kafka
Raymond Ops
Raymond Ops
Jun 28, 2026 · Databases

Comprehensive MySQL Replication Lag Troubleshooting Beyond Seconds_Behind_Master

This guide walks through a complete MySQL master‑slave lag diagnosis process, explaining why relying solely on Seconds_Behind_Master is insufficient and showing how to separate IO and SQL thread issues, examine relay logs, detect long transactions, DDL locks, and apply best‑practice configurations and monitoring.

LagMySQLReplication
0 likes · 17 min read
Comprehensive MySQL Replication Lag Troubleshooting Beyond Seconds_Behind_Master
MaGe Linux Operations
MaGe Linux Operations
Jun 28, 2026 · Operations

Practical Nginx Rate Limiting: Elegantly Defending Against CC Attacks and Traffic Spikes

This article walks through why Nginx needs rate limiting, explains the three core directives, compares burst, nodelay and delay behaviors, shows how to choose keys, and provides step‑by‑step configuration, testing, monitoring and troubleshooting recipes for protecting services from CC attacks and sudden traffic bursts.

Nginxcc-attacklimit_conn
0 likes · 29 min read
Practical Nginx Rate Limiting: Elegantly Defending Against CC Attacks and Traffic Spikes
Coder Trainee
Coder Trainee
Jun 27, 2026 · Backend Development

Mastering Java Thread‑Pool Tuning: Practical Performance Tips

This article explains why Java thread pools need tuning, walks through the seven core ThreadPoolExecutor parameters, provides formula‑based sizing, offers configuration templates for different workloads, shows monitoring and dynamic adjustment techniques, and highlights common pitfalls with concrete code examples.

JavaSpringThreadPoolExecutor
0 likes · 8 min read
Mastering Java Thread‑Pool Tuning: Practical Performance Tips
Raymond Ops
Raymond Ops
Jun 27, 2026 · Operations

Hands‑On DNS Ops: Deploy BIND and CoreDNS with Full Troubleshooting Guide

This comprehensive guide walks you through DNS fundamentals, compares BIND, CoreDNS, PowerDNS and Unbound, provides step‑by‑step deployment scripts for BIND 9.20 and CoreDNS 1.12, explains DNSSEC configuration, caching optimizations, security hardening, high‑availability designs, monitoring, backup and recovery procedures, and advanced troubleshooting techniques.

BINDCoreDNSDNS
0 likes · 43 min read
Hands‑On DNS Ops: Deploy BIND and CoreDNS with Full Troubleshooting Guide