Tagged articles

Operations

3429 articles · Page 1 of 35
Linyb Geek Road
Linyb Geek Road
Oct 6, 2026 · Operations

AI Writes Kubernetes YAML in Seconds: The Real Value of Ops Engineers

The article tests AI tools like DeepSeek for generating Kubernetes YAML, finding they handle standard templates well but fail on cluster-specific configs, security, probes, resource quotas, and complex multi-CRD scenarios. It argues ops engineers' value lies in troubleshooting, architecture decisions, incident handling, setting standards, and building platforms—not writing YAML—and advises embracing AI for drafts while deepening core expertise.

AIDevOpsKubernetes
0 likes · 13 min read
AI Writes Kubernetes YAML in Seconds: The Real Value of Ops Engineers
Golang Shines
Golang Shines
Sep 29, 2026 · Operations

Linux sysctl Tuning: Diagnose Bottlenecks Before Tweaking Kernel Parameters

A comprehensive guide to Linux kernel parameter tuning via sysctl, covering network, memory, and filesystem subsystems with a systematic workflow: baseline capture, bottleneck identification, single-parameter changes, validation, persistence, and rollback — illustrated with real production post-mortems.

LinuxNetworkingOperations
0 likes · 50 min read
Linux sysctl Tuning: Diagnose Bottlenecks Before Tweaking Kernel Parameters
Golang Shines
Golang Shines
Sep 26, 2026 · Operations

Complete Disk I/O Alert Troubleshooting: From Alert to Root Cause with 5 Real Cases

This article details a complete disk I/O alert investigation in production, covering core concepts like IOPS vs throughput, iostat/iotop analysis, and five real-world cases including MySQL missing indexes, log misconfiguration, backup conflicts, Redis persistence, and filesystem mount options, providing a reusable troubleshooting methodology.

MySQLOperationsPrometheus
0 likes · 65 min read
Complete Disk I/O Alert Troubleshooting: From Alert to Root Cause with 5 Real Cases
Efficient Ops
Efficient Ops
Sep 23, 2026 · Operations

Uptime Kuma Review: 77.7K Stars, Lightweight Self-Hosted Monitoring Setup

This article reviews Uptime Kuma, a lightweight open-source self-hosted monitoring tool with 77.7K GitHub stars, detailing its multi-protocol monitoring, 90+ notification channels, 20-second check intervals, and providing step-by-step installation guides for Docker Compose and manual Node.js deployments with PM2 process management.

DockerMonitoringNode.js
0 likes · 5 min read
Uptime Kuma Review: 77.7K Stars, Lightweight Self-Hosted Monitoring Setup
Random Bulletin
Random Bulletin
Sep 22, 2026 · Backend Development

Fault Prediction at 10M QPS: From Zero to Production-Ready System

This article details a practical roadmap for building production-grade fault prediction systems at massive scale, covering target selection, data governance, model evolution, time-window design, and safe action loops—emphasizing that reliable prediction requires engineering rigor beyond just model training.

Operationsfault predictionincident response
0 likes · 29 min read
Fault Prediction at 10M QPS: From Zero to Production-Ready System
Code Farmer Manor Chronicle
Code Farmer Manor Chronicle
Sep 22, 2026 · Operations

Windows Nginx Quick-Start: Install, Configure, and Run Without a VM

This guide covers downloading, extracting, and running Nginx on Windows, explaining directory structure differences from Linux, essential commands (start, reload, stop), configuration syntax with path handling, location matching priority, root vs alias behavior, and reverse proxy pitfalls like proxy_pass trailing slash and required headers.

NginxOperationsconfiguration
0 likes · 10 min read
Windows Nginx Quick-Start: Install, Configure, and Run Without a VM
Raymond Ops
Raymond Ops
Sep 21, 2026 · Operations

Nginx Rate Limiting in Practice: Defending Against CC Attacks and Traffic Spikes

This comprehensive guide covers Nginx rate limiting fundamentals, leaky bucket algorithm, configuration directives (limit_req_zone, limit_req, limit_conn_zone), practical scenarios (IP/URI-based limiting, whitelists, blacklists, CC attack defense), testing tools (wrk, ab, vegeta), monitoring with Prometheus, troubleshooting cases, and production rollout/rollback strategies.

NginxOperationsPrometheus
0 likes · 42 min read
Nginx Rate Limiting in Practice: Defending Against CC Attacks and Traffic Spikes
dbaplus Community
dbaplus Community
Sep 19, 2026 · Operations

AI Generates K8s YAML in Seconds: Where Is the Ops Engineer's Value?

The author tests AI tools like DeepSeek and ChatGPT for Kubernetes YAML generation, finding they handle standard templates well but fail on cluster-specific configs, security hardening, probe tuning, resource sizing, and complex multi-CRD scenarios, arguing ops value shifts from writing YAML to troubleshooting, architecture decisions, and platform building.

AIDevOpsKubernetes
0 likes · 13 min read
AI Generates K8s YAML in Seconds: Where Is the Ops Engineer's Value?
Golang Shines
Golang Shines
Sep 19, 2026 · Operations

SEAL Methodology for Production Troubleshooting: Veteran Ops Toolbox & Case Studies

A 10-year operations veteran shares the SEAL troubleshooting framework (Symptom, Environment, Analysis, Location), a curated toolbox (Prometheus, ELK, perf, tcpdump), real-world case studies (Redis avalanche, MySQL slow queries), incident grading, automation scripts, performance tuning, container/Kubernetes diagnostics, monitoring models, chaos engineering, and AIOps trends.

AIOpsMonitoringOperations
0 likes · 20 min read
SEAL Methodology for Production Troubleshooting: Veteran Ops Toolbox & Case Studies
Raymond Ops
Raymond Ops
Sep 18, 2026 · Operations

find vs locate: Mastering Million-File Search & Safe Cleanup in Linux Operations

This comprehensive guide compares find and locate for large-scale Linux file retrieval, covering performance bottlenecks, safe deletion practices, inode exhaustion troubleshooting, and production-ready scripts with dry-run, trash-based cleanup, and monitoring integration for million-file environments.

LinuxOperationsfile-search
0 likes · 76 min read
find vs locate: Mastering Million-File Search & Safe Cleanup in Linux Operations
Golang Shines
Golang Shines
Sep 17, 2026 · Operations

500 Essential Ops Terms: Kubernetes, Docker & SRE Glossary

This glossary defines 500 fundamental terms for operations engineers, covering Kubernetes core concepts, components, networking, and Docker container terminology with concise explanations for each term.

Container OrchestrationDevOpsDocker
0 likes · 5 min read
500 Essential Ops Terms: Kubernetes, Docker & SRE Glossary
Raymond Ops
Raymond Ops
Sep 15, 2026 · Operations

6 Battle-Tested Directions to Diagnose Nginx 502 Errors Fast

A systematic troubleshooting guide for Nginx 502 Bad Gateway errors covering six root-cause areas: upstream process/socket issues, config mismatches, permission blocks, application crashes/timeouts, DNS/TCP upstream problems, and host resource exhaustion — with exact commands, config snippets, and a ready-to-run evidence collection script.

502GunicornNginx
0 likes · 36 min read
6 Battle-Tested Directions to Diagnose Nginx 502 Errors Fast
Random Bulletin
Random Bulletin
Sep 15, 2026 · Operations

Structured Incident Response at 10M QPS: From Random to Process-Driven

This article details a comprehensive framework for transforming ad-hoc incident response into a structured, repeatable process for high-concurrency systems, covering incident state machines, role definitions, severity grading, first 15-minute checklists, timeline management, automation, blameless postmortems, and evolutionary stages from visibility to organizational learning.

High ConcurrencyOperationsautomation
0 likes · 30 min read
Structured Incident Response at 10M QPS: From Random to Process-Driven
Raymond Ops
Raymond Ops
Sep 14, 2026 · Operations

RocketMQ Production Operations: Cluster Setup, Retry Mechanisms & Dead Letter Queue Solutions

This comprehensive guide covers RocketMQ production operations including cluster deployment with NameServer and Broker configurations, message retry mechanisms with backoff strategies, dead letter queue handling and reprocessing, monitoring with Prometheus alerts, and troubleshooting procedures for common issues like message accumulation, disk full, and broker failures.

Cluster DeploymentDead Letter QueueMessage Queue
0 likes · 67 min read
RocketMQ Production Operations: Cluster Setup, Retry Mechanisms & Dead Letter Queue Solutions
Raymond Ops
Raymond Ops
Sep 11, 2026 · Operations

Master Linux Ops: Essential High-Frequency Commands for Daily Production Use

A comprehensive, scenario-driven reference covering 30+ categories of Linux operations commands — system info, processes, CPU, memory, network, disk, files, logs, users, services, packages, scheduling, performance analysis, text processing, SSH, troubleshooting workflows, dangerous commands, auditing, and efficiency tips — each with purpose, key parameters, real-world examples, and risk warnings.

DevOpsLinuxOperations
0 likes · 59 min read
Master Linux Ops: Essential High-Frequency Commands for Daily Production Use
Random Bulletin
Random Bulletin
Sep 11, 2026 · Operations

Complete Change Audit: Unified IDs, Identity Chains & Real-Time Risk Control

This article details how to evolve from basic operation logs to a complete change audit system for large-scale architectures by using unified change identifiers, identity chains, immutable evidence, and observability correlation to connect intent, authorization, execution, and impact for real-time risk control and incident reconstruction.

Operationschange auditchange management
0 likes · 40 min read
Complete Change Audit: Unified IDs, Identity Chains & Real-Time Risk Control
dbaplus Community
dbaplus Community
Sep 10, 2026 · Operations

Linux rm -rf Disaster Recovery: From Incident Response to File System Forensics

A comprehensive guide to recovering from accidental rm -rf deletions on Linux, covering immediate response, file system internals (ext4/xfs/btrfs/zfs), multiple recovery tools (extundelete, debugfs, lsof, snapshots), a real incident timeline, risk assessment, and prevention strategies including code review, monitoring, and backup practices.

EXT4LVMLinux
0 likes · 81 min read
Linux rm -rf Disaster Recovery: From Incident Response to File System Forensics
Cloud Architecture
Cloud Architecture
Sep 10, 2026 · Operations

From Firefighting to Fire Prevention: Production-Grade Database Monitoring with Prometheus & Grafana

This comprehensive guide details building a production-grade database monitoring system using Prometheus and Grafana, covering SLI/SLO design, alerting strategies, architecture, metric selection, security, Prometheus configuration, Alertmanager routing, Grafana dashboards, scaling, incident response runbooks, anti-patterns, and operational processes to shift from reactive firefighting to proactive prevention.

AlertingAlertmanagerDatabase Monitoring
0 likes · 42 min read
From Firefighting to Fire Prevention: Production-Grade Database Monitoring with Prometheus & Grafana
DevOps Operations Practice
DevOps Operations Practice
Sep 9, 2026 · Operations

Deploy Kafka 3.8.1 KRaft Cluster in 10 Minutes with Docker Compose

This guide walks through deploying a three-node Kafka 3.8.1 cluster using KRaft mode and Docker Compose, covering machine preparation, kernel tuning, cluster ID generation, node-specific configuration, startup, and verification steps including topic creation and message production/consumption.

Cluster SetupDockerDocker Compose
0 likes · 9 min read
Deploy Kafka 3.8.1 KRaft Cluster in 10 Minutes with Docker Compose
Raymond Ops
Raymond Ops
Sep 9, 2026 · Operations

Linux Logging Deep Dive: Kernel, journald, rsyslog & 5 Real Fault Cases

This comprehensive guide dissects the Linux logging stack — kernel ring buffer, journald, rsyslog, logrotate, and service logs — with configuration details, command references, and five step-by-step troubleshooting cases covering SSH brute force, disk exhaustion, OOM kills, network packet loss, and systemd service failures.

LinuxOperationsTroubleshooting
0 likes · 47 min read
Linux Logging Deep Dive: Kernel, journald, rsyslog & 5 Real Fault Cases
MaGe Linux Operations
MaGe Linux Operations
Sep 5, 2026 · Operations

Troubleshooting High Redis Client Connections: A Step-by-Step Guide to Identification and Optimization

This comprehensive guide details a systematic approach to diagnosing and resolving high Redis client connection counts, covering connection models, configuration tuning, CLI analysis commands, application-level connection pool fixes, temporary mitigation tactics, and long-term monitoring best practices with real-world case examples.

CLIENT LISTJedisMonitoring
0 likes · 35 min read
Troubleshooting High Redis Client Connections: A Step-by-Step Guide to Identification and Optimization
Architecture Digest
Architecture Digest
Sep 4, 2026 · Operations

Ongrid: Open-Source AI Agent Automates Full-Cycle Incident Response

The article reviews Ongrid, an open-source AI operations agent that automates alert investigation by querying metrics, logs, and traces, maps service topology for impact analysis, supports multiple LLMs, enforces read-only actions with approval gates, manages Kubernetes clusters, includes a built-in monitoring stack, workflow orchestration, knowledge base, and skill catalog, and provides installation steps and use cases.

AI AgentKubernetesOngrid
0 likes · 11 min read
Ongrid: Open-Source AI Agent Automates Full-Cycle Incident Response
Raymond Ops
Raymond Ops
Sep 1, 2026 · Operations

Master Linux File Permissions: How to Use chmod and chown Effectively

This comprehensive guide explains Linux's permission model, demonstrates numeric and symbolic chmod usage, details chown operations, introduces ACL for fine‑grained control, and provides troubleshooting steps and security best practices for production environments.

ACLLinuxOperations
0 likes · 33 min read
Master Linux File Permissions: How to Use chmod and chown Effectively
Random Bulletin
Random Bulletin
Sep 1, 2026 · Operations

From Manual to Automatic: Scaling Alert Automation for Million‑QPS Systems

The article examines why manual alert handling stalls at massive scale, outlines the risks of naïve auto‑rollback, and presents a step‑by‑step framework—including event control planes, executable runbooks, safety guards, and staged automation—to reliably move from human‑only to fully automated incident response in high‑throughput environments.

OperationsRunbookalert automation
0 likes · 23 min read
From Manual to Automatic: Scaling Alert Automation for Million‑QPS Systems
IT Architects Alliance
IT Architects Alliance
Sep 1, 2026 · Artificial Intelligence

Why Unlimited Token Budgets Let Agents Waste Money on Retries

The article explains how AI agents that automatically retry and invoke tools can silently accumulate hidden costs, argues for splitting billing into four categories, binding budgets to individual tasks, and handling over‑budget situations to prevent runaway expenses and long‑term maintenance burdens.

AI AgentsCost ManagementOperations
0 likes · 6 min read
Why Unlimited Token Budgets Let Agents Waste Money on Retries
Random Bulletin
Random Bulletin
Aug 30, 2026 · Operations

Alert Tiering: From a Single Level to a P0‑P3 Multi‑Level Response System

The article examines why a single‑level alerting approach fails at massive scale, outlines the five practical limitations it creates, and presents a step‑by‑step framework—classification matrix, routed channels, SLA timers, escalation policies, and on‑call discipline—to allocate limited human attention to the most business‑critical incidents.

AlertingMonitoringOperations
0 likes · 18 min read
Alert Tiering: From a Single Level to a P0‑P3 Multi‑Level Response System
Java Companion
Java Companion
Aug 28, 2026 · Operations

Beszel: A Sub‑10 MB Open‑Source Monitoring Panel That Can Replace Prometheus

Beszel is a lightweight, open‑source server‑monitoring dashboard written in Go that runs a Hub‑Agent architecture, offers per‑machine history charts, Docker container stats, alerting, multi‑user access and backup, and can be deployed in about ten minutes with Docker Compose, making it ideal for small‑scale environments where Prometheus feels heavyweight.

AlertingBeszelDocker Compose
0 likes · 9 min read
Beszel: A Sub‑10 MB Open‑Source Monitoring Panel That Can Replace Prometheus
Ops Community
Ops Community
Aug 26, 2026 · Operations

How to Diagnose Intermittent Packet Loss Using mtr, ss, and tcpdump

This guide walks you through a systematic, layer‑by‑layer approach to pinpointing occasional network packet loss by examining application logs, checking TCP connection stats with ss, reviewing kernel counters, tracing routes with mtr, and analyzing traffic with tcpdump, then applying targeted fixes.

LinuxOperationsmtr
0 likes · 32 min read
How to Diagnose Intermittent Packet Loss Using mtr, ss, and tcpdump
dbaplus Community
dbaplus Community
Aug 25, 2026 · Operations

What a Bank IT Leader Learned in 200 Days Replacing VMware

A mid‑size bank’s IT infrastructure head recounts a 200‑day journey swapping VMware for domestic virtualization, detailing performance gaps, CPU overcommit limits, memory management, backup reliability, compliance hurdles, and the step‑by‑step migration strategy that balanced risk and cost.

OperationsVMwarebanking
0 likes · 11 min read
What a Bank IT Leader Learned in 200 Days Replacing VMware
CTO Full-Stack Academy
CTO Full-Stack Academy
Aug 22, 2026 · Industry Insights

A Complete Blueprint for Implementing Smart Parks: Definitions, Architecture, and Step‑by‑Step Consulting Guide

This article provides a comprehensive, consulting‑driven roadmap for building smart parks, covering core definitions, a four‑layer cloud‑edge‑device architecture, a seven‑phase implementation process, common pitfalls with mitigation strategies, and a detailed case study that quantifies ROI and operational benefits.

IoTOperationsROI
0 likes · 26 min read
A Complete Blueprint for Implementing Smart Parks: Definitions, Architecture, and Step‑by‑Step Consulting Guide
Top Architecture Tech Stack
Top Architecture Tech Stack
Aug 21, 2026 · Operations

Why a ChatGPT and Codex Outage Shows You Need a Backup Model

The recent ChatGPT and Codex outage reveals that relying on a single AI entry point can cripple development pipelines, so teams should adopt layered fault handling, status tagging, externalized context, circuit‑breakers, human‑approved actions, and a lightweight backup model to maintain continuity.

AI reliabilityChatGPTCodex
0 likes · 9 min read
Why a ChatGPT and Codex Outage Shows You Need a Backup Model
Efficient Ops
Efficient Ops
Aug 19, 2026 · Operations

8 Must-Have MCP Ops Components That Dramatically Boost Efficiency

The article introduces eight essential MCP components—Grafana, Jenkins, K8s, Playwright, GitHub, Zabbix, Prometheus, and Alibaba Cloud—detailing how each enhances monitoring, automation, resource management, and performance optimization to cut fault‑resolution time, lower manual effort, and improve system stability.

GrafanaKubernetesMCP
0 likes · 7 min read
8 Must-Have MCP Ops Components That Dramatically Boost Efficiency
Old Zhao – Management Systems Only
Old Zhao – Management Systems Only
Aug 19, 2026 · Operations

10 Mindless Procurement Habits (And Why the Third Is the Worst)

The article reveals ten common mind‑less procurement habits—such as chasing low prices without total cost analysis, relying on memory instead of data, and ignoring inventory alerts—illustrates real‑world examples, and shows how turning these tasks into data‑driven processes with a digital platform can transform procurement from reactive fire‑fighting to proactive, strategic management.

Operationscost analysisdata-driven
0 likes · 11 min read
10 Mindless Procurement Habits (And Why the Third Is the Worst)
Linux Tech Enthusiast
Linux Tech Enthusiast
Aug 18, 2026 · Operations

Production Incident Troubleshooting Framework and Toolbox: Veteran Ops Engineer’s Real‑World Tips

A seasoned operations veteran shares a step‑by‑step incident‑response workflow, the SEAL troubleshooting methodology, essential monitoring and debugging tools, real‑world case studies, automated scripts, and best‑practice guidelines to help engineers quickly diagnose and resolve production outages.

LinuxMonitoringOperations
0 likes · 16 min read
Production Incident Troubleshooting Framework and Toolbox: Veteran Ops Engineer’s Real‑World Tips
Random Bulletin
Random Bulletin
Aug 16, 2026 · Operations

Real‑Time Alerting at Million‑QPS: From Static Thresholds to Intelligent Detection

At massive scales of millions of QPS and hundreds of metrics, static alert thresholds become noisy and hard to maintain; the article walks through a stepwise evolution—adding sustained‑duration and multi‑condition rules, adopting dynamic baselines, leveraging anomaly detection, and applying correlation, RCA and SLO burn‑rate techniques—to transform alerts from simple threshold breaches into precise, user‑experience‑focused notifications while combating alert fatigue.

AIOpsAlertingMonitoring
0 likes · 18 min read
Real‑Time Alerting at Million‑QPS: From Static Thresholds to Intelligent Detection
Linux Tech Enthusiast
Linux Tech Enthusiast
Aug 16, 2026 · Operations

Free Online Diagram Tools Every Ops Engineer Should Use

The article presents six free online diagramming platforms—Excalidraw, Zen Flowchart, Visual Paradigm Online, draw.io, 迅捷画图, and ProcessOn—detailing their key features, collaboration capabilities, template libraries, and direct URLs, helping operations professionals quickly choose the right visual‑communication tool.

ExcalidrawOperationscollaboration
0 likes · 7 min read
Free Online Diagram Tools Every Ops Engineer Should Use
Cloud Architecture
Cloud Architecture
Aug 15, 2026 · Cloud Native

Kubernetes Certificate Expiration Demystified: Incident Postmortem & 11‑Step Renewal Guide

The article analyzes a production outage caused by expired Kubernetes control‑plane certificates, explains why the failure cascades across components, and provides a detailed 11‑step procedure—including backup, certificate checks, etcd recovery, rolling restarts, and long‑term governance—to safely renew certificates in kubeadm‑based multi‑master clusters.

KubernetesOperationsautomation
0 likes · 37 min read
Kubernetes Certificate Expiration Demystified: Incident Postmortem & 11‑Step Renewal Guide
Linyb Geek Road
Linyb Geek Road
Aug 15, 2026 · Operations

Key Metrics Every Ops Engineer Should Monitor

This article enumerates essential operational metrics—such as CPU, memory, disk and network I/O, response time, throughput, error rates, availability, MTBF/MTTR, security logs, and capacity‑planning indicators—explaining their meanings and recommended target values to help engineers comprehensively monitor system performance, stability, and efficiency.

MonitoringOperationsavailability
0 likes · 10 min read
Key Metrics Every Ops Engineer Should Monitor
Old Zhao – Management Systems Only
Old Zhao – Management Systems Only
Aug 10, 2026 · Operations

How to Build Effective Supply Chain Planning: The Three‑Layer Framework and Six Key Actions

Many managers wonder why their planning teams still face material shortages, high inventory, and chaotic capacity, and the article explains that fragmented, unstructured plans are to blame, then introduces a three‑layer planning framework with six concrete actions and provides a ready‑to‑use template.

OperationsProcess OptimizationSupply Chain
0 likes · 2 min read
How to Build Effective Supply Chain Planning: The Three‑Layer Framework and Six Key Actions
Golang Shines
Golang Shines
Aug 8, 2026 · Interview Experience

How to Master 100+ Top Tech Ops Interview Questions and Land 20k+ Offers

This article compiles over 100 common DevOps and operations interview questions sourced from leading Chinese tech firms such as ByteDance, Meituan, Alibaba, Tencent and JD, providing a comprehensive study guide for candidates aiming to secure high‑salary offers.

CI/CDDevOpsInfrastructure as Code
0 likes · 4 min read
How to Master 100+ Top Tech Ops Interview Questions and Land 20k+ Offers
Coder Trainee
Coder Trainee
Aug 5, 2026 · Operations

How an Overnight Billing System Saved a Logistics Firm from Losing 3000 Yuan Daily

A logistics company struggled with mismatched label fees, manual reconciliation errors, and monthly profit loss, so the team deployed an automated recharge, label import, one‑click deduction, and reconciliation system that eliminated leakage, cut reconciliation time, and provided real‑time balance visibility.

MySQLOperationsSpring Boot
0 likes · 5 min read
How an Overnight Billing System Saved a Logistics Firm from Losing 3000 Yuan Daily
liandk
liandk
Aug 5, 2026 · Databases

Mastering Redis High Availability: Replication, Sentinel, and Cluster Explained

The article explains why Redis must be highly available and walks through three progressive architectures—master‑slave replication, Sentinel automatic failover, and Redis Cluster—detailing their mechanisms, advantages, drawbacks, and when to choose each for small, medium, or large‑scale production systems.

ClusterDatabase ScalingOperations
0 likes · 7 min read
Mastering Redis High Availability: Replication, Sentinel, and Cluster Explained
Ops Community
Ops Community
Aug 5, 2026 · Operations

Linux Kernel Sysctl Tuning Checklist – Proven Steps to Improve Performance

This article debunks the myth that simply copying a sysctl.conf yields a 30% boost, and presents a rigorous engineering loop—baseline measurement, hypothesis formulation, gray‑scale changes, observation of side effects, and rollback—along with detailed scripts, metrics, and per‑parameter guidance for memory, network, file handles, and more.

LinuxMonitoringOperations
0 likes · 37 min read
Linux Kernel Sysctl Tuning Checklist – Proven Steps to Improve Performance
Raymond Ops
Raymond Ops
Aug 4, 2026 · Operations

Uncover Hidden Nginx 502 Bad Gateway Config Pitfalls from Logs

This article explains why 502 Bad Gateway errors are the most frequent Nginx issue, quantifies their impact on business availability, and provides a systematic, log‑driven troubleshooting workflow with concrete configuration examples, health‑check setups, and production‑grade best‑practice recommendations.

502NginxOperations
0 likes · 73 min read
Uncover Hidden Nginx 502 Bad Gateway Config Pitfalls from Logs
Data Party THU
Data Party THU
Aug 4, 2026 · Operations

Why Multi-Agent Systems Are Fundamentally Distributed Systems

The article argues that multi‑agent workflows behave like traditional distributed systems, showing how deadlocks, state pollution, and silent drift arise from coordination failures rather than AI shortcomings, and it offers concrete engineering practices—timeouts, idempotency, cycle detection, and audit trails—to build reliable production‑grade agent pipelines.

Operationsdeadlockdistributed systems
0 likes · 14 min read
Why Multi-Agent Systems Are Fundamentally Distributed Systems
Golang Shines
Golang Shines
Aug 3, 2026 · Cloud Native

How I Built a Production‑Ready HA Kubernetes Cluster in Minutes

When my manager suddenly demanded a production‑grade, highly available Kubernetes cluster integrated with a private Harbor registry, I followed a comprehensive step‑by‑step guide to finish the entire setup within a few hours, and now share the 83‑page manual for anyone to replicate.

Cluster DeploymentHarborKubernetes
0 likes · 3 min read
How I Built a Production‑Ready HA Kubernetes Cluster in Minutes
Linyb Geek Road
Linyb Geek Road
Aug 2, 2026 · Operations

What Makes This Ops Expert’s Monitoring System Design So Effective?

The article explains how to build a comprehensive monitoring system using the USE method, outlines essential system and application metrics, and walks through the architecture and components of Prometheus, Grafana, full‑link tracing, and the ELK stack for effective operations monitoring.

ELKMonitoringOperations
0 likes · 13 min read
What Makes This Ops Expert’s Monitoring System Design So Effective?
Golang Shines
Golang Shines
Aug 1, 2026 · Operations

When a Snapshot Leak Triggered a P0 Outage: Lessons on Manual Cloud Ops

A hurried snapshot‑sharing command set public=true, unintentionally exposing customer data to all tenants, leading to a panic‑filled rollback, a painful post‑mortem, and a series of hard‑earned lessons about avoiding manual high‑risk operations, enforcing audit controls, and demanding productized UI for cloud infrastructure tasks.

Cloud ComputingOperationsRisk Management
0 likes · 10 min read
When a Snapshot Leak Triggered a P0 Outage: Lessons on Manual Cloud Ops
DataFunSummit
DataFunSummit
Jul 31, 2026 · Operations

Why Observability Agents Still Can’t Confirm Root Causes Despite Wider Connectors

Grafana Assistant now queries over 30 data sources, expanding incident clues across monitoring, databases, and ticket systems, but cross‑source access only improves correlation; without unified entity mapping, time alignment, and evidence verification, engineers cannot reliably prove a root cause.

Cross-Source QueryGrafana AssistantOperations
0 likes · 12 min read
Why Observability Agents Still Can’t Confirm Root Causes Despite Wider Connectors
AI Engineering
AI Engineering
Jul 30, 2026 · Operations

MCP’s Biggest Update: Stateless Core Eliminates Sessions, Handshakes, and Lowers Remote Deployment Barriers

The latest MCP release replaces the stateful session model with a stateless request/response core, removing initialize/initialized handshakes, enabling any server instance to handle requests via round‑robin, and adding features like MRTR, header‑based routing, cacheable list results, stronger auth, task extensions, and deprecations, which dramatically simplify operations and scaling.

MCPMRTROperations
0 likes · 7 min read
MCP’s Biggest Update: Stateless Core Eliminates Sessions, Handshakes, and Lowers Remote Deployment Barriers
DevOps Operations Practice
DevOps Operations Practice
Jul 30, 2026 · Operations

Essential Velero Guide for Kubernetes Disaster Recovery

This article walks through using Velero to back up, restore, and migrate Kubernetes clusters, covering MinIO installation, Velero client and server setup, storage volume creation, backup location configuration, and execution of backup, restore, and scheduled backup commands.

KubernetesOperationsVelero
0 likes · 11 min read
Essential Velero Guide for Kubernetes Disaster Recovery
dbaplus Community
dbaplus Community
Jul 28, 2026 · Operations

PostgreSQL Ops Checklist: 13 Pitfalls to Avoid (2026 Edition)

The article provides a comprehensive PostgreSQL operational self‑checklist covering CPU, memory, generic plans, idle connections, autovacuum, checkpoint and bgwriter tuning, AI workload handling, index building, replication, migration, partition management, and observability, with concrete examples, benchmarks, and practical commands.

OperationsPostgreSQLReplication
0 likes · 39 min read
PostgreSQL Ops Checklist: 13 Pitfalls to Avoid (2026 Edition)
Random Bulletin
Random Bulletin
Jul 28, 2026 · Operations

From Wiki Docs to Executable Incident Plans: Cutting MTTR from Hours to Minutes

The article explains how to evolve static wiki‑based incident response plans into structured, executable, auditable systems—adding one‑click execution, gray‑scale rollbacks, permission controls, chaos‑engineered rehearsals, and alert integration—to reduce mean‑time‑to‑recovery from hours to minutes in high‑throughput environments.

MTTR reductionOperationschaos engineering
0 likes · 19 min read
From Wiki Docs to Executable Incident Plans: Cutting MTTR from Hours to Minutes
Old Zhao – Management Systems Only
Old Zhao – Management Systems Only
Jul 27, 2026 · Operations

What Do the Eight Core Supply‑Chain Systems (ERP, WMS, MES, APS, TMS, S&OP/IBP, SRM, BI) Actually Do?

The article walks through the eight major supply‑chain management systems—ERP, WMS, MES, APS, TMS, S&OP/IBP, SRM, and BI—explaining the specific problems each solves, where they fit in a real‑world supply‑chain flow, and why simply implementing all of them does not guarantee a smooth supply chain.

ERPMESOperations
0 likes · 2 min read
What Do the Eight Core Supply‑Chain Systems (ERP, WMS, MES, APS, TMS, S&OP/IBP, SRM, BI) Actually Do?
PMTalk Product Manager Community
PMTalk Product Manager Community
Jul 27, 2026 · Artificial Intelligence

Why Your AI Skill Falls Short and How to Refine It in Three Real‑World Scenarios

The article explains why many AI Skills are unreliable, identifies three common failure patterns, and provides concrete scenario‑based refinements for product managers, designers, and operators, along with practical checklists and management tips to turn a draft Skill into a stable, reusable workflow.

AIDesign ReviewOperations
0 likes · 19 min read
Why Your AI Skill Falls Short and How to Refine It in Three Real‑World Scenarios
ITPUB
ITPUB
Jul 26, 2026 · Operations

When Cutting Ops Staff Breaks the System: Real Cost of Layoffs

Multiple real‑world anecdotes show that eliminating operations personnel—whether senior SQL optimizers, script‑maintaining engineers, or on‑site ops staff—triggers hidden expenses, system outages, and massive productivity loss that far outweigh any short‑term savings.

Cost ManagementIT staffingOperations
0 likes · 7 min read
When Cutting Ops Staff Breaks the System: Real Cost of Layoffs
Linyb Geek Road
Linyb Geek Road
Jul 26, 2026 · Operations

From Alert Flood to Fault Insight: The Real Starting Point of AIOps

The article explains that successful AIOps begins not with sophisticated models but with turning a flood of fragmented alerts into a single, context‑rich incident view that tells operators how many failures occurred, which business services are impacted, and where they should start investigating.

AIOpsMonitoringOperations
0 likes · 14 min read
From Alert Flood to Fault Insight: The Real Starting Point of AIOps
Linyb Geek Road
Linyb Geek Road
Jul 26, 2026 · Operations

Postmortem: How an Alert Flood Masked the Real Problem

A late‑night incident flooded the on‑call channel with dozens of red alerts, hiding the true root cause—a core service latency spike—until the team re‑ordered information, prioritized early signals, and applied a simple three‑tier alert classification to restore clarity and speed up resolution.

AIOpsMonitoringOperations
0 likes · 11 min read
Postmortem: How an Alert Flood Masked the Real Problem
DeepHub IMBA
DeepHub IMBA
Jul 24, 2026 · Operations

Avoid Repeating Microservice Governance Pitfalls in AI Agent Management

The article analyzes how AI agents create hidden, "shadow" integrations that are harder to detect than traditional services, outlines five critical governance questions, and proposes a set of operational capabilities and principles—identity, observability, governance, lifecycle, and reuse—to responsibly scale AgentOps.

AI AgentOperationsShadow Integration
0 likes · 10 min read
Avoid Repeating Microservice Governance Pitfalls in AI Agent Management
Tech Architecture Stories
Tech Architecture Stories
Jul 22, 2026 · Operations

Running a 17‑Week Automated WeChat Publishing Workflow with WorkBuddy

For 17 consecutive weeks, WorkBuddy automatically triggers at 9 am every Sunday, fetches the top GitHub projects, generates a markdown report, syncs a website candidate pool, creates a WeChat draft, and sends a result email, all built with a system‑design approach that ensures fault isolation, state management, and repeatable execution.

AI workflowContent systemGitHub automation
0 likes · 21 min read
Running a 17‑Week Automated WeChat Publishing Workflow with WorkBuddy
Ops Community
Ops Community
Jul 21, 2026 · Operations

How to Extend Zabbix Without Writing Any Code

This article presents a code‑free Zabbix Agent deployment module that lets administrators batch‑install agents via the Zabbix web UI, explains its key features, typical use cases, step‑by‑step installation instructions, and showcases the resulting monitoring setup with screenshots.

Agent DeploymentCSV ImportCustom Module
0 likes · 7 min read
How to Extend Zabbix Without Writing Any Code
Golang Shines
Golang Shines
Jul 18, 2026 · Operations

149 Essential Shell Scripts for Sysadmins – Keep Them Handy

This article provides a curated collection of 149 practical Bash scripts for Linux system administrators, covering common tasks such as locating zombie processes, removing empty files, calculating sums and ranges, formatting dates, retrieving MAC addresses, checking leap years, sorting numbers, and more, with each example presented as ready‑to‑run code.

LinuxOperationsautomation
0 likes · 8 min read
149 Essential Shell Scripts for Sysadmins – Keep Them Handy
Efficient Ops
Efficient Ops
Jul 16, 2026 · Operations

Why Harness, Not Model, Is the Real Key for Deploying AI Agents in Production

After a 40‑minute GOPS talk on "one person + an AI Agent army," the author argues that the decisive factor for putting AI agents into production is not the model itself but a disciplined harness engineering that creates a verifiable, roll‑backable closed loop, turning a single operator into a minimal delivery unit.

AI AgentsOPCOperations
0 likes · 13 min read
Why Harness, Not Model, Is the Real Key for Deploying AI Agents in Production
Code of Duty
Code of Duty
Jul 15, 2026 · Operations

From Idea to ICP Filing: How I Built the Entry Layer of My Personal Blog

The article reviews the first phase of creating a personal blog, explaining why a blog serves as a tech‑brand foundation, how server, domain, and ICP filing form the essential entry layer, and what pitfalls to avoid before moving to deployment and security.

ICP filingOperationsdomain registration
0 likes · 11 min read
From Idea to ICP Filing: How I Built the Entry Layer of My Personal Blog
MaGe Linux Operations
MaGe Linux Operations
Jul 15, 2026 · Operations

Essential Git Commands Every Ops Engineer Needs for Reliable Deployments

This guide walks operations engineers through the core Git concepts, status checks, cloning strategies, branch handling, release directory design, conflict resolution, worktree usage, bisect debugging, and safe push practices, providing concrete command examples and scripts to make deployments reproducible and auditable.

CI/CDGitOperations
0 likes · 30 min read
Essential Git Commands Every Ops Engineer Needs for Reliable Deployments
Raymond Ops
Raymond Ops
Jul 14, 2026 · Operations

Network Troubleshooting with tcpdump & Wireshark: Step‑by‑Step Guide and Ready‑to‑Use Scripts

This comprehensive guide walks you through using tcpdump and Wireshark for network fault isolation, covering core concepts, capture filters, detailed analysis techniques, performance tuning, expert information interpretation, automation scripts, and best‑practice recommendations for efficient packet‑level troubleshooting.

LinuxOperationsWireshark
0 likes · 51 min read
Network Troubleshooting with tcpdump & Wireshark: Step‑by‑Step Guide and Ready‑to‑Use Scripts
Efficient Ops
Efficient Ops
Jul 13, 2026 · Operations

Open‑Source Nginx UI: A Visual Tool That Can Triple Ops Efficiency

Nginx UI offers a graphical interface for configuring and monitoring Nginx, includes real‑time metrics, extensible modules and AI Agent integration, and provides multiple installation options such as systemd, Docker and a one‑click script, promising up to three‑fold productivity gains for operators.

AI AgentDockerNginx
0 likes · 5 min read
Open‑Source Nginx UI: A Visual Tool That Can Triple Ops Efficiency
Efficient Ops
Efficient Ops
Jul 12, 2026 · Operations

How China Telecom’s Dual Systems Earned SRE Level‑3 Certification and Elevated SOMM Operations

China Telecom’s production‑grade CPCP and enterprise‑marketing systems successfully passed the CAICT SRE Level‑3 assessment, achieving over 99.9% annual availability, zero incidents, and significant improvements in observability, chaos engineering, automation, and capacity planning, as detailed in an interview with senior IT managers.

AIOpsChina TelecomDevOps
0 likes · 14 min read
How China Telecom’s Dual Systems Earned SRE Level‑3 Certification and Elevated SOMM Operations
Raymond Ops
Raymond Ops
Jul 8, 2026 · Operations

10 Essential System Commands Every Ops Engineer Should Master

This guide explains why core Linux commands are indispensable for ops engineers, categorizes them into monitoring, networking, disk analysis, text processing and service management, and provides detailed usage examples, troubleshooting scenarios, and decision‑tree guidance for effective system administration.

LinuxMonitoringOperations
0 likes · 47 min read
10 Essential System Commands Every Ops Engineer Should Master
Ops Community
Ops Community
Jul 5, 2026 · Operations

20 Common Ops Newbie Pitfalls – Which Ones Have You Hit?

This guide catalogs the 20 most frequent mistakes made by new operations engineers, explains why they happen, and provides step‑by‑step safe alternatives, risk warnings, and recovery procedures so readers can avoid costly outages and build reliable habits.

DevOpsKubernetesLinux
0 likes · 29 min read
20 Common Ops Newbie Pitfalls – Which Ones Have You Hit?
MaGe Linux Operations
MaGe Linux Operations
Jul 4, 2026 · Operations

20 Common Ops Rookie Mistakes and How to Avoid Them

This guide lists the twenty most frequent pitfalls that new operations engineers encounter, explains why they happen, and provides step‑by‑step safe practices, code examples, risk classifications and a verification checklist to help prevent costly outages and data loss.

DevOpsKubernetesLinux
0 likes · 28 min read
20 Common Ops Rookie Mistakes and How to Avoid Them
Raymond Ops
Raymond Ops
Jul 3, 2026 · Operations

10 Rookie Ops Mistakes You Must Avoid – A Complete Checklist

This guide walks ops newcomers through the ten most common pitfalls—from accidental rm‑rf deletions and mis‑configured firewalls to unsafe chmod usage—and provides concrete remediation steps, ready‑to‑run shell scripts, best‑practice checklists, and monitoring setups to keep production environments stable and secure.

DevOpsLinuxMonitoring
0 likes · 51 min read
10 Rookie Ops Mistakes You Must Avoid – A Complete Checklist
Architect
Architect
Jul 1, 2026 · Artificial Intelligence

Scheduling AI Agents for Night‑Shift Work: Turning Prompts into Reliable Loops

The article explains how to transform AI agents from single‑prompt responders into reliable night‑shift workers by defining clear goals, state files, evidence, and permission boundaries, using /goal, /loop and scheduled tasks, and provides concrete steps, examples, and a scheduling template for stable unattended execution.

AI AgentsGoalLoop
0 likes · 27 min read
Scheduling AI Agents for Night‑Shift Work: Turning Prompts into Reliable Loops
Tencent Cloud Developer
Tencent Cloud Developer
Jul 1, 2026 · Fundamentals

What Is Architecture Really? Business, Application, and Data Views

The article explores the true meaning of architecture by distinguishing business, application, data, and technical architectures, explains how architecture consists of elements, structure, and connections, and provides practical guidelines, common pitfalls, design principles, and evolution paths from monolithic to distributed and micro‑service systems.

OperationsR&D Managementbackend development
0 likes · 23 min read
What Is Architecture Really? Business, Application, and Data Views
Random Bulletin
Random Bulletin
Jun 29, 2026 · Operations

Backlog Digestion at Ten‑Million QPS: From Adding Machines to Intelligent Scheduling

The article dissects how to handle massive message backlog in ten‑million‑QPS systems, explaining why simply adding consumer machines fails, and walks through six evolutionary stages—from manual scaling and auto‑scaling to traffic tiering, consumer‑side optimizations, dynamic strategies, and AI‑driven intelligent scheduling—while highlighting design trade‑offs, pitfalls, and practical tooling.

Message QueueOperationsauto-scaling
0 likes · 22 min read
Backlog Digestion at Ten‑Million QPS: From Adding Machines to Intelligent Scheduling
Smart Workplace Lab
Smart Workplace Lab
Jun 29, 2026 · Operations

Why 20‑person Group Chats Stall for Days and How Frontline Owners Can Use an Asynchronous Consensus Convergence SOP

The article analyzes why large asynchronous group discussions waste up to 72 hours without a decision, introduces a three‑step “divergence extraction + delegated decision + execution verification” protocol that cuts convergence time to 12 hours (‑80 %), reduces manual effort by 75 % and can be deployed in ten minutes using built‑in AI and approval tools.

AI AutomationOperationsProcess Optimization
0 likes · 7 min read
Why 20‑person Group Chats Stall for Days and How Frontline Owners Can Use an Asynchronous Consensus Convergence SOP
Frontend AI Walk
Frontend AI Walk
Jun 29, 2026 · Operations

Loop Engineering: Which Scenarios Really Work and Which to Avoid

The article defines three screening criteria—repetition, verifiability, and worth—to evaluate Loop Engineering tasks, lists six high‑value scenarios ranging from code engineering to business operations, warns against unsuitable use cases, and provides a step‑by‑step onboarding guide.

AI AgentsLoop EngineeringOperations
0 likes · 12 min read
Loop Engineering: Which Scenarios Really Work and Which to Avoid
ITPUB
ITPUB
Jun 28, 2026 · Industry Insights

What’s the Longest‑Running Server Ever? Real‑World Uptime Stories

The article compiles dozens of real‑world examples of computers and servers that have stayed online for years or even decades, from a Chinese provincial telecom data‑center Red Hat Linux box running 14 years, to a 20‑year‑old base‑station, a 20‑plus‑year DOS server, two Linux boxes up since 2007, and NASA’s Voyager 2 spacecraft computer that has been operating for over 43 years.

DOSOperationsRed Hat Linux
0 likes · 8 min read
What’s the Longest‑Running Server Ever? Real‑World Uptime Stories
TechVision Expert Circle
TechVision Expert Circle
Jun 26, 2026 · Operations

How CTOs Can Build Systems That Make Their Own Decisions

The article explains why, in 2026, CTOs must equip production systems with self‑decision capabilities, outlines an OODA‑loop‑based architecture with perception, decision (three‑brain LLM agent), execution, and feedback layers, and addresses practical challenges such as latency, hallucinations, cost, and team adoption.

LLMOODAOperations
0 likes · 14 min read
How CTOs Can Build Systems That Make Their Own Decisions
Smart Workplace Lab
Smart Workplace Lab
Jun 25, 2026 · Operations

Cut Approval Time by 80% with a Single Excel Sheet—No IT Changes Needed

The article outlines a step‑by‑step, Excel‑based workflow that identifies approval bottlenecks, creates a group whitelist, and implements lightweight SOPs to shave up to 80% off approval cycle time, saving two hours daily and letting teams focus on high‑risk items without requiring system changes.

ExcelOperationsSOP
0 likes · 7 min read
Cut Approval Time by 80% with a Single Excel Sheet—No IT Changes Needed