Tagged articles

Runbook

5 articles · Page 1 of 1
Cloud Architecture
Cloud Architecture
Sep 10, 2026 · Operations

From Firefighting to Fire Prevention: Production-Grade Database Monitoring with Prometheus & Grafana

This comprehensive guide details building a production-grade database monitoring system using Prometheus and Grafana, covering SLI/SLO design, alerting strategies, architecture, metric selection, security, Prometheus configuration, Alertmanager routing, Grafana dashboards, scaling, incident response runbooks, anti-patterns, and operational processes to shift from reactive firefighting to proactive prevention.

AlertmanagerDatabase MonitoringGrafana
0 likes · 42 min read
From Firefighting to Fire Prevention: Production-Grade Database Monitoring with Prometheus & Grafana
Random Bulletin
Random Bulletin
Sep 1, 2026 · Operations

From Manual to Automatic: Scaling Alert Automation for Million‑QPS Systems

The article examines why manual alert handling stalls at massive scale, outlines the risks of naïve auto‑rollback, and presents a step‑by‑step framework—including event control planes, executable runbooks, safety guards, and staged automation—to reliably move from human‑only to fully automated incident response in high‑throughput environments.

ObservabilityRunbookalert automation
0 likes · 23 min read
From Manual to Automatic: Scaling Alert Automation for Million‑QPS Systems
Raymond Ops
Raymond Ops
Apr 20, 2026 · Operations

How to Build a Standardized SRE On‑Call Process: From Alert Grading to Handoff Templates

This article presents a complete SRE on‑call handbook that defines alert severity levels, provides concrete Prometheus Alertmanager configurations, outlines a step‑by‑step response flow, details war‑room roles, escalation paths, handoff checklists, post‑mortem procedures, and dozens of ready‑to‑use templates to reduce MTTR and improve reliability.

RunbookSREalert management
0 likes · 27 min read
How to Build a Standardized SRE On‑Call Process: From Alert Grading to Handoff Templates
MaGe Linux Operations
MaGe Linux Operations
Dec 10, 2025 · Operations

Standardized SRE On‑Call Handbook: Alert Grading, Response Flow, and Handoff Templates

This handbook presents a complete, two‑year‑tested SRE on‑call process that defines alert severity tiers, response requirements, escalation paths, War‑Room roles, handoff schedules, post‑mortem procedures, and provides ready‑to‑use configuration snippets, checklists and templates to reduce MTTR and repeat incidents.

RunbookSREalert management
0 likes · 26 min read
Standardized SRE On‑Call Handbook: Alert Grading, Response Flow, and Handoff Templates
Meituan Technology Team
Meituan Technology Team
Apr 14, 2017 · Operations

Meituan Dianping's Automation Concepts and Practices

Meituan Dianping’s technical salon showcased its automation journey—detailing database automation platforms, service‑tree management, Puppet‑based web control, and a CMDB case study from Shanghai Zhaogang—highlighting rapid iteration, standardization challenges, and the evolution of automation from tools to core operational practice.

Cloud ComputingDatabase AutomationDevOps
0 likes · 4 min read
Meituan Dianping's Automation Concepts and Practices