Tagged articles

SRE

431 articles · Page 5 of 5
Efficient Ops
Efficient Ops
Aug 29, 2017 · Operations

From ITIL to SRE: How Vipshop Transformed Its Operations

This article recounts Vipshop’s journey from a traditional ITIL‑based operations model to an SRE‑inspired, automated workflow, detailing the construction of ITIL processes, the challenges faced, the shift toward automation, and personal insights on managing people, quality, and change.

DevOpsITILSRE
0 likes · 20 min read
From ITIL to SRE: How Vipshop Transformed Its Operations
Efficient Ops
Efficient Ops
Aug 28, 2017 · Operations

Can Ops Teams Become Agile? A Practical Kanban Journey

This article explores how operations teams can adopt agile principles—especially Kanban—to address common challenges such as delayed feedback, task overload, and hidden risks, demonstrating a step‑by‑step transformation within the DevOps lifecycle.

AgileDevOpsKanban
0 likes · 28 min read
Can Ops Teams Become Agile? A Practical Kanban Journey
Efficient Ops
Efficient Ops
Jul 25, 2017 · Operations

Why Google’s SRE Model Matters: Lessons for Modern Ops Teams

This article explains the origins, responsibilities, and team structures of Google Site Reliability Engineering (SRE), compares it with traditional operations roles in companies like Yahoo, Alibaba, and Facebook, and offers practical guidance for building effective SRE or application‑operations teams today.

DevOpsSREsite reliability engineering
0 likes · 25 min read
Why Google’s SRE Model Matters: Lessons for Modern Ops Teams
Efficient Ops
Efficient Ops
Jun 10, 2017 · Operations

What Google’s SRE Book Reveals About Modern Operations

This article introduces the Chinese translation of Google’s SRE book, shares behind‑the‑scenes stories of its creation, and distills key concepts such as the AAA model, Borg architecture, SLOs, toil reduction, and the cultural shift required for reliable large‑scale services.

DevOpsGoogleInfrastructure
0 likes · 20 min read
What Google’s SRE Book Reveals About Modern Operations
ITPUB
ITPUB
Jun 9, 2017 · Operations

Mastering Effective Monitoring: From Basics to the USE Method

This article explains the fundamentals of monitoring, distinguishes traditional OPS from SRE perspectives, defines monitoring objects and metrics, introduces quantitative thinking with SLI/SLO, and presents the USE method with a MySQL example to help engineers detect and prevent failures efficiently.

SLISLOSRE
0 likes · 10 min read
Mastering Effective Monitoring: From Basics to the USE Method
Efficient Ops
Efficient Ops
May 25, 2017 · Operations

How a Bank Transformed IT Ops with Automated DevOps and SRE Practices

This article outlines how China Merchants Bank’s data‑center application management team identified traditional financial IT operational pain points, introduced DevOps and SRE concepts, built non‑functional management frameworks, and implemented automated tooling, monitoring, and capacity‑scaling to achieve fully automated operations.

DevOpsIT OperationsPerformance Scaling
0 likes · 24 min read
How a Bank Transformed IT Ops with Automated DevOps and SRE Practices
ITPUB
ITPUB
May 15, 2017 · Operations

Mastering Online Incident Management: From Detection to Prevention

This article outlines a comprehensive methodology for handling large‑scale online service incidents, covering goals, the "jump‑fill‑avoid" framework, step‑by‑step processes for detection, diagnosis, remediation, and post‑mortem analysis, as well as essential monitoring, logging, and escalation infrastructure.

SREincident managementlog tracing
0 likes · 18 min read
Mastering Online Incident Management: From Detection to Prevention
360 Zhihui Cloud Developer
360 Zhihui Cloud Developer
Apr 6, 2017 · Operations

How SRE’s Dialectical Thinking Redefines Modern Operations

An insightful reflection on Google’s SRE philosophy shows how dialectical thinking—questioning absolute stability, embracing limited toil, prioritizing simple monitoring, recognizing automation’s hidden risks, and practicing real‑world failure drills—can reshape operations, encouraging smarter, more resilient system design.

SREautomationmonitoring
0 likes · 7 min read
How SRE’s Dialectical Thinking Redefines Modern Operations
Efficient Ops
Efficient Ops
Mar 26, 2017 · Operations

How Google Scales App Engine: Lessons in Cloud Scalability and SRE

The article shares Google SRE veteran Minghua Ye’s insights on App Engine’s evolution, emphasizing the critical role of automatic scalability, distributed locks, service discovery, load balancing, and open‑source tools like gRPC, Protobuf, gflags, glog, and Googletest in building reliable, high‑traffic cloud services.

Google App EngineProtobufSRE
0 likes · 12 min read
How Google Scales App Engine: Lessons in Cloud Scalability and SRE
Efficient Ops
Efficient Ops
Mar 21, 2017 · Operations

Rethinking Operations: The “Third Kind” of SRE at Lianjia

The article shares the author’s experience transitioning from private to public and hybrid clouds at Lianjia, introduces a “third kind” of operations that blends traditional and internet‑based practices, and discusses containers, DNS‑based naming, and automation tools to build adaptable, cost‑effective infrastructure.

InfrastructureNaming ServiceSRE
0 likes · 21 min read
Rethinking Operations: The “Third Kind” of SRE at Lianjia
High Availability Architecture
High Availability Architecture
Mar 15, 2017 · Operations

Highlights from SRECon17 Americas 2023 in San Francisco

The article reports on the SRECon17 Americas conference in San Francisco, summarizing keynote talks, panel sessions, and practical insights from industry leaders such as Stripe, Netflix, Google, and IBM on topics ranging from traffic control and container management to on‑call practices and cost considerations for Site Reliability Engineering.

DevOpsGoogleNetflix
0 likes · 6 min read
Highlights from SRECon17 Americas 2023 in San Francisco
Ctrip Technology
Ctrip Technology
Dec 9, 2016 · Operations

Design and Implementation of Ctrip Call Center's Active‑Active Architecture and Unified Login

The article details Ctrip's call‑center architecture evolution, describing the multi‑layer active‑active design, public access, application and client layers, unified login mechanisms, operational challenges, disaster‑recovery drills, and future plans for software‑only and mobile agents, illustrating practical SRE principles in a large‑scale telephony system.

Active-ActiveIP phoneSRE
0 likes · 22 min read
Design and Implementation of Ctrip Call Center's Active‑Active Architecture and Unified Login
Efficient Ops
Efficient Ops
Dec 4, 2016 · Operations

How Ctrip Built a Seamless Multi‑Region Dual‑Active Call Center

This article details Ctrip's evolution from a single‑site call‑center to a fully dual‑active, multi‑region architecture, covering the overall system design, public network, application, and client layers, unified login mechanisms, heartbeat monitoring, and future software‑only and mobile‑first directions.

Dual-ActiveSREcall center
0 likes · 27 min read
How Ctrip Built a Seamless Multi‑Region Dual‑Active Call Center
Efficient Ops
Efficient Ops
Nov 7, 2016 · Operations

How to Train New SREs Effectively: Proven Practices and Playbooks

This article outlines a systematic approach to onboarding and training new Site Reliability Engineers, covering trust building, readiness assessment, diverse learning methods, structured curricula, on‑call milestones, project‑focused work, reverse‑engineering skills, statistical thinking, and improvisation techniques to develop high‑performing SRE teams.

SREon-calloperations
0 likes · 17 min read
How to Train New SREs Effectively: Proven Practices and Playbooks
Efficient Ops
Efficient Ops
Oct 30, 2016 · Operations

How Google Music Recovered 1.5 PB of Lost Data After a Massive Deletion Bug

In March 2012, a privacy‑driven deletion pipeline mistakenly erased hundreds of thousands of Google Music files, prompting SREs to launch a massive data‑recovery effort that involved MapReduce impact analysis, tape‑based backups, and a complete redesign of the deletion system.

Google MusicLarge-Scale DeletionSRE
0 likes · 14 min read
How Google Music Recovered 1.5 PB of Lost Data After a Massive Deletion Bug
Efficient Ops
Efficient Ops
Oct 26, 2016 · Operations

From Sysadmin to Google SRE: How Modern Ops Teams Can Thrive

This article compares traditional system administration with Google’s Site Reliability Engineering, explaining why enterprises are shifting from cost‑center SLA focus to data‑driven, user‑experience‑oriented operations, and offers practical steps for teams to adopt automation, cloud platforms, and risk‑aware practices.

SRE
0 likes · 14 min read
From Sysadmin to Google SRE: How Modern Ops Teams Can Thrive
Efficient Ops
Efficient Ops
Oct 23, 2016 · Operations

How Google’s SRE Postmortems Drive System Reliability

This article explains Google’s SRE postmortem philosophy, the criteria for writing postmortems, best practices for a blame‑free culture, and how collaborative knowledge‑sharing and incentives improve incident handling and overall system reliability.

SREincident managementoperations
0 likes · 14 min read
How Google’s SRE Postmortems Drive System Reliability
Efficient Ops
Efficient Ops
Oct 5, 2016 · Operations

5 Must‑Read Ops ‘Prescriptions’ to Boost Your Infrastructure Skills

This article curates five technical reads—covering network operations, Google’s production environment, massive cost‑saving strategies, IDC automation, and Docker‑based RDS—each presented as a “medicine” with a brief description and a link for deeper insight.

Cloud ComputingCost OptimizationDocker
0 likes · 5 min read
5 Must‑Read Ops ‘Prescriptions’ to Boost Your Infrastructure Skills
Efficient Ops
Efficient Ops
Sep 18, 2016 · Operations

Who Was the World’s First SRE? Uncovering Margaret Hamilton’s Legacy

This article explores the origins of Site Reliability Engineering, highlights Margaret Hamilton as the likely first SRE through her work on NASA’s Apollo program, and draws lessons on reliability, disaster prevention, and the evolution of modern SRE practices.

Apollo programMargaret HamiltonSRE
0 likes · 10 min read
Who Was the World’s First SRE? Uncovering Margaret Hamilton’s Legacy
Efficient Ops
Efficient Ops
Sep 13, 2016 · Operations

How Google SRE Principles Compare Across Industries

This article, excerpted from the upcoming Chinese edition of “SRE: Google Site Reliability Engineering”, examines how Google’s SRE guiding philosophies—disaster planning, post‑mortem culture, automation, and data‑driven decision‑making—are adopted, adapted, or contrasted in sectors such as manufacturing, aerospace, nuclear, telecommunications, healthcare, and finance, highlighting key similarities, differences, and lessons for Google and the broader tech industry.

SREautomationincident response
0 likes · 21 min read
How Google SRE Principles Compare Across Industries
MaGe Linux Operations
MaGe Linux Operations
May 27, 2016 · Operations

Why Google Relies on Software Engineers to Run Its Services: Inside SRE

The article explains Google’s Site Reliability Engineering (SRE) philosophy, how it empowers software engineers to automate operations, the balance between development and reliability, the concept of error budgets, and the cultural shift that turned DevOps into a core practice for large‑scale services.

DevOpsError BudgetSRE
0 likes · 10 min read
Why Google Relies on Software Engineers to Run Its Services: Inside SRE
Efficient Ops
Efficient Ops
Jul 27, 2015 · Operations

What Google SREs Do: Inside the Role that Powers Reliable Services

This article explains the responsibilities, requirements, and daily work of Google Site Reliability Engineers, contrasts them with Software Engineers, outlines key internal infrastructure components, and discusses the future direction of operations engineering in the cloud era.

GoogleInfrastructureSRE
0 likes · 11 min read
What Google SREs Do: Inside the Role that Powers Reliable Services
MaGe Linux Operations
MaGe Linux Operations
Apr 28, 2015 · Operations

How Yelp Achieved Zero‑Downtime HAProxy Reloads Using Linux qdisc

Yelp’s infrastructure team tackled HAProxy’s reload‑induced packet loss by leveraging Linux’s plug qdisc and iptables to delay SYN packets during reloads, enabling zero‑downtime service updates and improving reliability despite the kernel’s brief binding window.

HAProxyLinux qdiscNetwork Traffic Control
0 likes · 7 min read
How Yelp Achieved Zero‑Downtime HAProxy Reloads Using Linux qdisc