Tagged articles

Kubernetes

4438 articles · Page 1 of 45
Linyb Geek Road
Linyb Geek Road
Oct 6, 2026 · Operations

AI Writes Kubernetes YAML in Seconds: The Real Value of Ops Engineers

The article tests AI tools like DeepSeek for generating Kubernetes YAML, finding they handle standard templates well but fail on cluster-specific configs, security, probes, resource quotas, and complex multi-CRD scenarios. It argues ops engineers' value lies in troubleshooting, architecture decisions, incident handling, setting standards, and building platforms—not writing YAML—and advises embracing AI for drafts while deepening core expertise.

AICloud NativeDevOps
0 likes · 13 min read
AI Writes Kubernetes YAML in Seconds: The Real Value of Ops Engineers
Random Bulletin
Random Bulletin
Oct 1, 2026 · Backend Development

Fault Domain Design: Turning Blast Radius from 100% into a Tunable 1/N Parameter

The article presents a layered fault-domain strategy—physical anti-affinity, logical isolation (sharding, cluster groups, swimlanes, bulkheads), cell-based architecture, chaos-engineering validation, and quantitative governance metrics—to shrink the blast radius of a ten-million-QPS system from a fixed 100% to a controllable 1/N design parameter.

KubernetesSLOanti-affinity
0 likes · 25 min read
Fault Domain Design: Turning Blast Radius from 100% into a Tunable 1/N Parameter
Ops Development & AI Practice
Ops Development & AI Practice
Sep 25, 2026 · Cloud Native

Apache APISIX Evolution: From Disrupting Kong to Premier AI Gateway

This article traces Apache APISIX's four-phase evolution: its etcd-based architecture achieving millisecond config updates versus Kong's seconds, its 9-month Apache graduation, multi-language plugin runner and Wasm support, Kubernetes-native ingress controller, and transformation into an AI Gateway with token-based rate limiting, protocol normalization, and SSE streaming optimization.

AI GatewayAPI GatewayApache APISIX
0 likes · 17 min read
Apache APISIX Evolution: From Disrupting Kong to Premier AI Gateway
Tencent Cloud Developer
Tencent Cloud Developer
Sep 22, 2026 · Backend Development

Changing the Engine Mid-Flight: Isolated Verification Environments for Login System Migration

A team replaced a core login system in two months without halting feature development by building isolated verification environments using Makefile-standardized builds, a lightweight Kubernetes platform, and custom debugging tools, reducing verification prep time from 30–60 minutes to ~5 minutes and enabling safe fault-injection testing.

AI-assisted developmentKubernetesMakefile
0 likes · 33 min read
Changing the Engine Mid-Flight: Isolated Verification Environments for Login System Migration
dbaplus Community
dbaplus Community
Sep 19, 2026 · Operations

AI Generates K8s YAML in Seconds: Where Is the Ops Engineer's Value?

The author tests AI tools like DeepSeek and ChatGPT for Kubernetes YAML generation, finding they handle standard templates well but fail on cluster-specific configs, security hardening, probe tuning, resource sizing, and complex multi-CRD scenarios, arguing ops value shifts from writing YAML to troubleshooting, architecture decisions, and platform building.

AIDevOpsKubernetes
0 likes · 13 min read
AI Generates K8s YAML in Seconds: Where Is the Ops Engineer's Value?
Ops Development & AI Practice
Ops Development & AI Practice
Sep 17, 2026 · Cloud Native

Why OpenTelemetry Helm Splits into 3 Releases: Agent, Cluster, Gateway Architecture Explained

This article explains why OpenTelemetry Helm charts now recommend deploying Collector as three separate releases—otel-agent (DaemonSet for node metrics), otel-cluster (singleton Deployment for cluster metrics), and otel-gateway (scalable Deployment for trace ingestion)—detailing Presets simplification, lifecycle isolation, failure domains, and when to consolidate to two releases.

Cloud NativeCollectorDaemonSet
0 likes · 24 min read
Why OpenTelemetry Helm Splits into 3 Releases: Agent, Cluster, Gateway Architecture Explained
Golang Shines
Golang Shines
Sep 17, 2026 · Operations

500 Essential Ops Terms: Kubernetes, Docker & SRE Glossary

This glossary defines 500 fundamental terms for operations engineers, covering Kubernetes core concepts, components, networking, and Docker container terminology with concise explanations for each term.

Cloud NativeContainer OrchestrationDevOps
0 likes · 5 min read
500 Essential Ops Terms: Kubernetes, Docker & SRE Glossary
YiSu Grain
YiSu Grain
Sep 16, 2026 · Fundamentals

12 Core Computer Science Problems Solved: CPU, Paging, Normalization & Kubernetes Probes

This article provides detailed solutions and explanations for 12 computer science practice problems covering CPU execution time calculation, bus bandwidth, virtual memory paging, subnet addressing, database normalization (2NF/3NF), lossless join decomposition, UML class diagram notation, Kubernetes probe types, earned value management, deadlock prevention strategies, and common architectural misconceptions.

KubernetesNetworkingOperating Systems
0 likes · 17 min read
12 Core Computer Science Problems Solved: CPU, Paging, Normalization & Kubernetes Probes
AI Digital Ideal
AI Digital Ideal
Sep 15, 2026 · Operations

DevOps Automation with DeepSeek Harness: Quality Gates, GitOps & Self-Healing

This chapter teaches DevOps automation practices using DeepSeek Harness, covering quality gates with AI code review, multi-environment deployment pipelines with blue-green strategy, GitOps workflows with ArgoCD and Kustomize, Terraform infrastructure as code, automated self-healing and cost optimization, and intelligent alert analysis for production operations.

Alert AnalysisCI/CDDeepSeek Harness
0 likes · 29 min read
DevOps Automation with DeepSeek Harness: Quality Gates, GitOps & Self-Healing
IT Services Circle
IT Services Circle
Sep 15, 2026 · Backend Development

How to Diagnose and Resolve CPU/Memory Spikes in Agent Batch Processing

This article details a systematic approach to diagnosing and resolving CPU and memory spikes when an Agent processes large document batches, covering backpressure propagation, concurrency budgeting, runtime profiling, and pipeline design to prevent OOM kills and throttling.

AgentBatch ProcessingCPU spike
0 likes · 23 min read
How to Diagnose and Resolve CPU/Memory Spikes in Agent Batch Processing
AI Digital Ideal
AI Digital Ideal
Sep 14, 2026 · Backend Development

Building an Enterprise Code Review Agent with DeepSeek Harness: Chapter 14 Practical Project

This chapter guides developers through constructing a production-ready code review agent using the DeepSeek Harness framework, covering architecture design, multi-analyzer integration (LSP, security, style, complexity), GitHub/GitLab webhook handling, automated reporting, Docker/Kubernetes deployment, and testing strategies.

DeepSeek HarnessDockerGitHub integration
0 likes · 31 min read
Building an Enterprise Code Review Agent with DeepSeek Harness: Chapter 14 Practical Project
Cloud Architecture
Cloud Architecture
Sep 13, 2026 · Backend Development

Production-Grade RAG with Spring AI: Verifiable, Rollbackable, Auditable Knowledge Base

This article details a production-ready customer service knowledge base built with Spring AI 2.0.1 and Milvus, covering immutable index versioning, tenant-isolated retrieval with parameterized filters, deterministic chunk IDs, idempotent ingestion pipelines, dual-index blue-green deployments, and comprehensive observability with automated rollback triggers.

KubernetesMilvusProduction Engineering
0 likes · 27 min read
Production-Grade RAG with Spring AI: Verifiable, Rollbackable, Auditable Knowledge Base
Architect's Guide
Architect's Guide
Sep 10, 2026 · Backend Development

Apollo Configuration Center: Complete Guide from Core Concepts to Kubernetes Deployment

This article provides a comprehensive tutorial on Apollo configuration center, covering its core concepts (application, environment, cluster, namespace), client architecture with long-polling and local caching, overall system design with Config/Admin services and Eureka, availability scenarios, hands-on SpringBoot integration with Maven setup, dynamic configuration testing (updates, rollbacks, offline fallback), multi-environment/cluster/namespace usage, and Kubernetes deployment via Docker and YAML manifests.

Configuration CenterCtripJava
0 likes · 27 min read
Apollo Configuration Center: Complete Guide from Core Concepts to Kubernetes Deployment
Full-Stack DevOps & Kubernetes
Full-Stack DevOps & Kubernetes
Sep 9, 2026 · Cloud Native

AI Low-Code Cuts K8s Platform Dev Time by 90%: A Practical Guide

The article demonstrates how AI low-code tools like Cursor accelerate building a Kubernetes management platform, reducing development from 7-10 days to 1-2 days, with code examples for Go/FastAPI backend, Vue3 frontend, plus cases for Prometheus alert forwarding and CRUD backends, while stressing human oversight for security and logic.

AI low-codeCloud NativeCursor
0 likes · 18 min read
AI Low-Code Cuts K8s Platform Dev Time by 90%: A Practical Guide
Code Mala Tang
Code Mala Tang
Sep 7, 2026 · Cloud Native

Kubernetes Resource Contracts: How requests and limits Control Pod Scheduling, Throttling, and Eviction

This article explains how Kubernetes requests and limits act as a contract between workloads and the scheduler, detailing their distinct roles, the divergent consequences of CPU throttling versus memory OOMKills, the three QoS classes that determine eviction priority, and how LimitRange and ResourceQuota enforce resource policies at the namespace level.

CPU throttlingKubernetesLimitRange
0 likes · 10 min read
Kubernetes Resource Contracts: How requests and limits Control Pod Scheduling, Throttling, and Eviction
Code Mala Tang
Code Mala Tang
Sep 6, 2026 · Cloud Native

Kubernetes Volume Mounts: Hot Reload, subPath Pitfalls & Projected Volumes

This article explains Kubernetes volume mounting for ConfigMaps and Secrets, detailing how directory mounts achieve hot reloads via atomic symlink switching, why subPath mounts never update, and how projected volumes merge multiple configuration sources into a single directory.

ConfigMapHot ReloadKubernetes
0 likes · 10 min read
Kubernetes Volume Mounts: Hot Reload, subPath Pitfalls & Projected Volumes
LuTiao Programming
LuTiao Programming
Sep 5, 2026 · Backend Development

Why @Scheduled Runs 3× in Kubernetes: Migrating to JobRunr for Distributed Tasks

After scaling a Spring Boot app to three Kubernetes replicas, @Scheduled tasks ran thrice, generating duplicate reports. The author evaluates distributed locks but adopts JobRunr for persistent, retryable, and monitorable background jobs, demonstrating integration with Spring Boot 4, code patterns for fire-and-forget, recurring, and delayed tasks, plus idempotency and deduplication strategies.

Background JobsDistributed Task SchedulingJobRunr
0 likes · 16 min read
Why @Scheduled Runs 3× in Kubernetes: Migrating to JobRunr for Distributed Tasks
Raymond Ops
Raymond Ops
Sep 5, 2026 · Operations

Linux Time Synchronization Mastery: Chrony Best Practices for Production Systems

Comprehensive guide covering Linux time concepts, NTP protocol, chrony vs ntpd, clock source selection, leap second handling, configuration templates for cloud, containers, Kubernetes, and isolated networks, plus verification, monitoring, troubleshooting, compliance automation, and rollback strategies.

KubernetesNTPTroubleshooting
0 likes · 54 min read
Linux Time Synchronization Mastery: Chrony Best Practices for Production Systems
Architecture Digest
Architecture Digest
Sep 4, 2026 · Operations

Ongrid: Open-Source AI Agent Automates Full-Cycle Incident Response

The article reviews Ongrid, an open-source AI operations agent that automates alert investigation by querying metrics, logs, and traces, maps service topology for impact analysis, supports multiple LLMs, enforces read-only actions with approval gates, manages Kubernetes clusters, includes a built-in monitoring stack, workflow orchestration, knowledge base, and skill catalog, and provides installation steps and use cases.

AI AgentKubernetesOngrid
0 likes · 11 min read
Ongrid: Open-Source AI Agent Automates Full-Cycle Incident Response
Raymond Ops
Raymond Ops
Sep 2, 2026 · Operations

Essential Network Basics for Ops: IP, Subnet Mask, and Gateway Explained in One Go

This guide walks operations engineers through IP addressing, binary‑decimal conversion, classful and private address ranges, subnet masks and CIDR notation, gateway functions, VLAN concepts, routing tables, DNS basics, Docker/Kubernetes networking, firewall rules, and practical troubleshooting steps, all illustrated with concrete commands and examples.

DNSDockerIP
0 likes · 36 min read
Essential Network Basics for Ops: IP, Subnet Mask, and Gateway Explained in One Go
Woodpecker Software Testing
Woodpecker Software Testing
Sep 2, 2026 · Operations

5 Common Pitfalls in Performance Regression Testing

In today’s fast‑paced agile and micro‑service environments, performance regression testing is often treated as optional, leading to severe TPS drops and hidden degradations; this article details five typical misconceptions, backs them with real‑world examples, and offers concrete practices to make performance regression a continuous, cross‑team responsibility.

CI/CDKubernetesLoad Testing
0 likes · 8 min read
5 Common Pitfalls in Performance Regression Testing
Woodpecker Software Testing
Woodpecker Software Testing
Sep 1, 2026 · Cloud Native

Why 90% of Container Performance Issues Come From Poor Capacity Planning – An In‑Depth Look

The article explains how container performance testing must evolve from simple load simulation to chaos‑engineered, observability‑driven capacity planning, introduces a 4‑dimensional capacity model, and shows automated SLI‑based scaling using real‑world e‑commerce and finance case studies.

Control PlaneKubernetescapacity planning
0 likes · 7 min read
Why 90% of Container Performance Issues Come From Poor Capacity Planning – An In‑Depth Look
dbaplus Community
dbaplus Community
Aug 31, 2026 · Interview Experience

Why Kubernetes Leader Kelsey Hightower Refused Microsoft and Retired at 43

From dropping out of college to becoming Google’s distinguished L9 engineer, Kelsey Hightower’s unconventional journey—spanning DSL modem work, A+ certification, data‑center roles, Puppet, CoreOS, and Kubernetes—reveals his belief that technology serves people, his refusal of Microsoft, and his early retirement at 43.

AICareerCloud Native
0 likes · 50 min read
Why Kubernetes Leader Kelsey Hightower Refused Microsoft and Retired at 43
Tencent Technical Engineering
Tencent Technical Engineering
Aug 31, 2026 · Cloud Native

CubeSandbox v0.7.0: Cross-Machine Sandbox Migration & Multi-Version Runtime Support

CubeSandbox v0.7.0 introduces cross-machine pause/resume via S3 backend storage, allows component multi-version coexistence so upgrades don't break existing templates, merges NetworkAgent into Cubelet to cut RPC calls and speed cold starts, and separates control-plane scheduling from node operations via new CubeOps service.

Cloud NativeCubeSandboxKubernetes
0 likes · 10 min read
CubeSandbox v0.7.0: Cross-Machine Sandbox Migration & Multi-Version Runtime Support
TonyBai
TonyBai
Aug 30, 2026 · Cloud Native

Stop Struggling with Kubernetes Docs: An Amazon‑Warehouse Analogy That Reveals the Whole System

Using an Amazon‑warehouse analogy, the article explains Kubernetes’s core philosophy of declarative desired state and maps each control‑plane and node component—apiserver, etcd, controller‑manager, scheduler, kubelet, kube‑proxy—to intuitive warehouse roles, while detailing service discovery, EndpointSlices, and external traffic flow.

Cloud NativeControl PlaneDesired State
0 likes · 17 min read
Stop Struggling with Kubernetes Docs: An Amazon‑Warehouse Analogy That Reveals the Whole System
Alibaba Cloud Native
Alibaba Cloud Native
Aug 29, 2026 · Cloud Native

Ingress NGINX Retired Amid New Critical Vulnerabilities – Migrate to Alibaba Cloud API Gateway in 10 Minutes

Ingress NGINX has been retired and is plagued by multiple CVSS 8.1 high‑severity vulnerabilities that lack patches, prompting urgent migration to Alibaba Cloud's Cloud Native API Gateway, which now offers expanded CLB/NLB reuse, annotation compatibility analysis, and integrated traffic‑shifting and rollback workflows.

ACKAPI GatewayCloud Native
0 likes · 14 min read
Ingress NGINX Retired Amid New Critical Vulnerabilities – Migrate to Alibaba Cloud API Gateway in 10 Minutes
AI Engineering
AI Engineering
Aug 28, 2026 · Cloud Native

Why Round‑Robin Fails for LLM Inference and How llm‑d Fixes It

Round‑robin routing in Kubernetes wipes out KV‑cache benefits for LLM inference, but llm‑d introduces cache‑aware routing, hierarchical eviction, and prefill/decode separation, delivering up to three‑fold throughput gains and halving first‑token latency, as shown in Tesla's production rollout.

Cloud NativeKV CacheKubernetes
0 likes · 7 min read
Why Round‑Robin Fails for LLM Inference and How llm‑d Fixes It
mikechen
mikechen
Aug 26, 2026 · Cloud Native

Kubernetes Architecture Deep Dive: Master-Node Components & Workflow

This article explains Kubernetes architecture, covering the master-node model, core components like API server, etcd, scheduler, controller manager, kubelet, and kube-proxy, and how they coordinate to deploy and manage containerized applications.

API ServerCloud NativeContainer Orchestration
0 likes · 6 min read
Kubernetes Architecture Deep Dive: Master-Node Components & Workflow
Cloud Architecture
Cloud Architecture
Aug 25, 2026 · Backend Development

Designing Nginx for Million‑Scale WebSocket Connections: Architecture, Configuration, and Pitfalls

This article walks through the end‑to‑end design of a production‑grade Nginx‑based WebSocket gateway that can handle a million concurrent connections, covering the five essential requirements, detailed Nginx settings, backend gateway responsibilities, load‑balancing strategies, observability, common failure patterns, and step‑by‑step Go code examples.

GoKubernetesNginx
0 likes · 34 min read
Designing Nginx for Million‑Scale WebSocket Connections: Architecture, Configuration, and Pitfalls
Yumin Fish Harvest
Yumin Fish Harvest
Aug 23, 2026 · Backend Development

How to Deploy TSID for Distributed ID Generation: From Local Setup to Multi‑Cluster Kubernetes

This article walks through using the TSID library to generate 64‑bit numeric and 13‑character string IDs, shows a quick‑start Maven example, integrates the creator into Spring Boot as a singleton bean, explains node and nodeCount configuration, and details deployment strategies for fixed servers, a single Kubernetes cluster, and multiple clusters.

JavaKubernetesSpring Boot
0 likes · 16 min read
How to Deploy TSID for Distributed ID Generation: From Local Setup to Multi‑Cluster Kubernetes
Ops Community
Ops Community
Aug 22, 2026 · Operations

Five Overlooked Runtime Risks When Deploying Large Language Models on Kubernetes

Deploying large‑model inference services on Kubernetes can hide five critical runtime risks—such as premature traffic before model loading, GPU memory overflow, LivenessProbe mis‑kills, slow HPA scaling, and missing logs—that only surface under production load, leading to timeouts, crashes, and costly debugging.

AIHPAKubernetes
0 likes · 33 min read
Five Overlooked Runtime Risks When Deploying Large Language Models on Kubernetes
TonyBai
TonyBai
Aug 21, 2026 · Cloud Native

VictoriaMetrics' vlagent Hits 143k Logs/sec—How It Outperforms 8 Popular Log Collectors

A rigorous benchmark of nine Kubernetes log collectors under a 1‑core, 1 GiB limit shows VictoriaMetrics' vlagent achieving 143,000 lines per second—4.5× faster than Fluent Bit and 28× faster than Fluentd—while using the least CPU and memory, and exposing hidden correctness bugs in several competitors.

Fluent BitKubernetesLog Collection
0 likes · 17 min read
VictoriaMetrics' vlagent Hits 143k Logs/sec—How It Outperforms 8 Popular Log Collectors
Airbnb Technology Team
Airbnb Technology Team
Aug 20, 2026 · Cloud Native

How Airbnb Built a Scalable, Reliable Kubernetes Sidecar for Dynamic Configuration

The article explains Airbnb's Sitar‑agent sidecar architecture, detailing the end‑to‑end configuration distribution lifecycle, key design choices such as sidecar versus in‑process deployment, pull‑model optimizations, and the migration from Sparkey to SQLite for robust, multi‑language support at massive scale.

Cloud NativeKubernetesRocksDB
0 likes · 15 min read
How Airbnb Built a Scalable, Reliable Kubernetes Sidecar for Dynamic Configuration
Efficient Ops
Efficient Ops
Aug 19, 2026 · Operations

8 Must-Have MCP Ops Components That Dramatically Boost Efficiency

The article introduces eight essential MCP components—Grafana, Jenkins, K8s, Playwright, GitHub, Zabbix, Prometheus, and Alibaba Cloud—detailing how each enhances monitoring, automation, resource management, and performance optimization to cut fault‑resolution time, lower manual effort, and improve system stability.

GrafanaKubernetesMCP
0 likes · 7 min read
8 Must-Have MCP Ops Components That Dramatically Boost Efficiency
YiSu Grain
YiSu Grain
Aug 19, 2026 · Cloud Native

Day 60 Cloud‑Native Case Study: Service Governance, Reliable Messaging, and Observability

This Day 60 case study walks through a regional medical appointment platform that has been broken into micro‑services on a container cluster, asking you to select and justify service discovery, TCC/Saga, reliable messaging, Kubernetes, Service Mesh and observability measures, and to explain their benefits and trade‑offs.

Distributed TransactionsKubernetesReliable Messaging
0 likes · 36 min read
Day 60 Cloud‑Native Case Study: Service Governance, Reliable Messaging, and Observability
Cloud Architecture
Cloud Architecture
Aug 17, 2026 · Backend Development

Go Microservice Stability: Rate Limiting, Circuit Breaking, Degradation and K8s Production Architecture

The article walks through a real‑world traffic spike in an e‑commerce order service, explains why isolated techniques like rate limiting, circuit breaking or degradation are insufficient, and presents a complete, layered stability‑governance solution for Go microservices running on Kubernetes, complete with code, configuration, observability and testing guidance.

Circuit BreakerGoKubernetes
0 likes · 42 min read
Go Microservice Stability: Rate Limiting, Circuit Breaking, Degradation and K8s Production Architecture
Cloud Architecture
Cloud Architecture
Aug 17, 2026 · Backend Development

Comprehensive Guide to Building an Enterprise‑Grade Distributed ID System in Go

This article walks through the full design and production‑ready implementation of a Go‑based distributed ID service, comparing Snowflake and Leaf Segment algorithms, detailing a dual‑engine architecture, SDK caching, scaling on Kubernetes, observability, deployment, and performance testing for high‑throughput enterprise applications.

GoKubernetesLeaf Segment
0 likes · 36 min read
Comprehensive Guide to Building an Enterprise‑Grade Distributed ID System in Go
Cloud Architecture
Cloud Architecture
Aug 15, 2026 · Cloud Native

Kubernetes Certificate Expiration Demystified: Incident Postmortem & 11‑Step Renewal Guide

The article analyzes a production outage caused by expired Kubernetes control‑plane certificates, explains why the failure cascades across components, and provides a detailed 11‑step procedure—including backup, certificate checks, etcd recovery, rolling restarts, and long‑term governance—to safely renew certificates in kubeadm‑based multi‑master clusters.

KubernetesOperationsautomation
0 likes · 37 min read
Kubernetes Certificate Expiration Demystified: Incident Postmortem & 11‑Step Renewal Guide
Cloud Architecture
Cloud Architecture
Aug 13, 2026 · Cloud Native

Kubernetes Node Maintenance: From Drain to True Zero‑Downtime Engineering

Many teams mistakenly believe that a simple `kubectl drain` guarantees safe node shutdown, but in production the risk spans the control plane, service discovery, long‑lived connections, load balancers and observability; this guide presents a repeatable, auditable, production‑grade process that turns node maintenance into a zero‑interruption engineering workflow.

Graceful ShutdownKubernetesOperator
0 likes · 41 min read
Kubernetes Node Maintenance: From Drain to True Zero‑Downtime Engineering
DataFunSummit
DataFunSummit
Aug 13, 2026 · Cloud Native

Agent Architecture Evolution: From Monolithic Self‑Management to Distributed Hosting

The article outlines a step‑by‑step evolution of Agent systems, explaining why traditional microservice patterns fail, describing three monolithic deployment models, detailing how separating session, memory, and environment state enables distributed hosting, and presenting function‑as‑a‑service to fully managed ReAct and multi‑Agent collaboration via Registry and A2A.

AgentAgent RegistryFunction-as-a-Service
0 likes · 14 min read
Agent Architecture Evolution: From Monolithic Self‑Management to Distributed Hosting
Random Bulletin
Random Bulletin
Aug 13, 2026 · Cloud Native

Canary Releases: From Simple Percentages to Precise, Attribute‑Based Deployments

The article examines how traditional percentage‑based canary releases evolve into fine‑grained, attribute‑driven deployments with session stickiness, automated SLO gating, traffic mirroring, and service‑mesh integration, highlighting five pain points, practical solutions, and the hidden prerequisites for large‑scale systems.

IstioKubernetesSLO
0 likes · 19 min read
Canary Releases: From Simple Percentages to Precise, Attribute‑Based Deployments
Geek Labs
Geek Labs
Aug 13, 2026 · Artificial Intelligence

How Centaur Enables a Secure, Unified Self‑Hosted AI Agent for the Whole Team

Centaur transforms personal AI coding assistants into a self‑hosted, team‑shared platform by deploying agents in isolated Kubernetes sandboxes, using iron‑proxy for credential injection, persisting workflows in Postgres, and providing Slack and HTTP interfaces, thus solving configuration duplication, credential leakage, context fragmentation, and audit challenges.

AI agentsKubernetessecurity
0 likes · 15 min read
How Centaur Enables a Secure, Unified Self‑Hosted AI Agent for the Whole Team
Java Architecture Diary
Java Architecture Diary
Aug 12, 2026 · Cloud Native

Why Upgrading Your MCP Server to 2.0 Solves Stateless Session Issues

The article explains how MCP 1.x's stateful handshake caused node‑crash failures, sticky sessions, and serverless incompatibility, and how the 2.0 release removes the handshake, makes each request self‑describing via _meta and HTTP headers, introduces MRTR for multi‑round interactions, and provides a Java/TypeScript code walkthrough demonstrating the new stateless behavior.

Cloud NativeJavaKubernetes
0 likes · 8 min read
Why Upgrading Your MCP Server to 2.0 Solves Stateless Session Issues
SpringMeng
SpringMeng
Aug 12, 2026 · Databases

RedisInsight: The Official High‑Performance GUI for Redis

This article introduces RedisInsight, the official visual management tool for Redis, outlines its key features, provides step‑by‑step installation on Linux and Kubernetes, and demonstrates basic usage for monitoring, querying, and memory analysis through the GUI.

GUIInstallationKubernetes
0 likes · 7 min read
RedisInsight: The Official High‑Performance GUI for Redis
Cloud Architecture
Cloud Architecture
Aug 11, 2026 · Databases

Redis Sentinel Deep Dive: Leader Election, Failover Mechanics, and Production Best Practices

This article dissects Redis Sentinel’s high‑availability workflow—from failure detection, SDOWN/ODOWN states, and quorum logic to leader election, replica promotion, and configuration propagation—while illustrating each step with a real‑world e‑commerce cache case, detailed configuration snippets, Kubernetes deployment patterns, Spring Boot integration, and operational playbooks for observability and fault‑injection testing.

KubernetesRedisSentinel
0 likes · 48 min read
Redis Sentinel Deep Dive: Leader Election, Failover Mechanics, and Production Best Practices
Full-Stack DevOps & Kubernetes
Full-Stack DevOps & Kubernetes
Aug 11, 2026 · Cloud Native

One‑Click AI Knowledge Base Solves Massive Document Search for K8s Fault Root‑Cause Analysis

The article describes a self‑built K8s‑RAG‑AIOps tool that uses an offline vector knowledge base and the DeepSeek‑v4‑pro model to automatically retrieve internal SOPs, collect live cluster data via SSH, and generate a complete, executable fault‑diagnosis report, dramatically speeding up Kubernetes troubleshooting while keeping data secure.

AIOpsDeepSeekDevOps
0 likes · 9 min read
One‑Click AI Knowledge Base Solves Massive Document Search for K8s Fault Root‑Cause Analysis
Cloud Architecture
Cloud Architecture
Aug 10, 2026 · Databases

Master‑Slave Replication in Redis: Core Mechanics Explained and Production Deployment

This article provides a comprehensive, production‑focused analysis of Redis master‑slave replication, covering its internal state machine, full and partial sync processes, configuration pitfalls, performance bottlenecks, consistency trade‑offs, and practical deployment patterns with Docker, Kubernetes, and Spring Boot.

Docker ComposeKubernetesRedis
0 likes · 38 min read
Master‑Slave Replication in Redis: Core Mechanics Explained and Production Deployment
Cloud Architecture
Cloud Architecture
Aug 10, 2026 · Cloud Native

How to Deploy Docker Images Offline Without Downtime: A Complete Enterprise Solution

This article presents a production‑grade, step‑by‑step solution for offline Docker image distribution in enterprise environments, covering OCI image fundamentals, layer reuse, digest‑based governance, a multi‑domain architecture with Harbor, Skopeo, Crane and Trivy, and practical scripts for building, exporting, validating, importing, and pre‑warming images across large Kubernetes clusters while ensuring security, compliance, and high‑concurrency performance.

CraneDockerHarbor
0 likes · 32 min read
How to Deploy Docker Images Offline Without Downtime: A Complete Enterprise Solution
Ray's Galactic Tech
Ray's Galactic Tech
Aug 10, 2026 · Cloud Native

Destruction and Rebirth: Deep Dive into ETCD Backup and Restore for Kubernetes Clusters

This article walks through a real‑world ETCD failure, explains why ETCD is the control‑plane brain, details the three‑layer ETCD architecture, exposes common backup pitfalls, and provides a production‑grade backup‑restore workflow—including snapshot API usage, Go implementation, verification steps, and post‑restore validation—for reliable Kubernetes disaster recovery.

Cloud NativeGoKubernetes
0 likes · 33 min read
Destruction and Rebirth: Deep Dive into ETCD Backup and Restore for Kubernetes Clusters
Golang Shines
Golang Shines
Aug 10, 2026 · Cloud Native

Build a Binary‑Based Kubernetes 1.36 Cluster from Scratch (PDF Guide)

This article explains how to manually assemble a Kubernetes 1.36 cluster using binary files, detailing the overall architecture, core components, and the roles of master and worker nodes, while providing a free 55‑page PDF with step‑by‑step instructions.

Binary InstallationControl PlaneKubernetes
0 likes · 3 min read
Build a Binary‑Based Kubernetes 1.36 Cluster from Scratch (PDF Guide)
Xiaolin Talks Programming
Xiaolin Talks Programming
Aug 10, 2026 · Big Data

Building Real-Time Data Lakes with Spring Boot, Flink CDC 3.0 & Iceberg: Production Patterns

This article details migrating from a legacy Canal+Kafka+Flink+Hive stack to a modern Flink CDC + Iceberg real-time data lake, covering Spring Boot orchestration, chunk-based full/incremental sync, Interval Join for wide tables, automated schema evolution, hidden partitioning, Exactly-Once checkpoint tuning, and Kubernetes deployment practices.

Apache IcebergExactly-OnceFlink CDC
0 likes · 17 min read
Building Real-Time Data Lakes with Spring Boot, Flink CDC 3.0 & Iceberg: Production Patterns
Ray's Galactic Tech
Ray's Galactic Tech
Aug 9, 2026 · Cloud Native

Goodbye Hand‑Written YAML Hell: Deploy Microservices with Helm

The article explains how manual Kubernetes YAML quickly becomes unmanageable in microservice environments and demonstrates how Helm provides templating, parameterization, versioning, dependency management, and lifecycle hooks to create a standardized, automated, production‑grade deployment pipeline that can handle billions of requests across multiple services.

CI/CDGitOpsHelm
0 likes · 35 min read
Goodbye Hand‑Written YAML Hell: Deploy Microservices with Helm
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Aug 9, 2026 · Cloud Native

How ACK One Fleet Transforms Agent Sandbox from Single-Cluster to Multi-Cluster

The article explains how ACK One Fleet upgrades the AI Agent Sandbox from a single‑cluster Kubernetes setup to a multi‑cluster architecture, addressing capacity limits, fault‑domain risks, and scheduling inefficiencies while providing global capacity control, water‑level balancing, fault‑tolerant failover, and faster sandbox startup through E2B and CRD integrations.

ACK OneCloud NativeE2B
0 likes · 11 min read
How ACK One Fleet Transforms Agent Sandbox from Single-Cluster to Multi-Cluster
Raymond Ops
Raymond Ops
Aug 8, 2026 · Operations

Mastering K8s Troubleshooting: Common Production Issues and Essential Commands

This guide walks you through the most frequent Kubernetes production problems—from pod failures like CrashLoopBackOff and ImagePullBackOff to node NotReady states, service DNS errors, storage PVC issues, RBAC permissions, and scheduling conflicts—providing step‑by‑step diagnostic commands, concrete examples, and practical remediation strategies to keep your clusters stable and your services running.

KubernetesNodeRBAC
0 likes · 51 min read
Mastering K8s Troubleshooting: Common Production Issues and Essential Commands
DataFunSummit
DataFunSummit
Aug 8, 2026 · Cloud Native

Building Containerized Sandboxes for Multi‑Agent AI: Architecture, Key Technologies, and Real‑World Practices

The article examines how to construct a container‑based sandbox infrastructure for multi‑agent AI systems, covering isolation mechanisms, lifecycle management, resource scaling, checkpoint/commit techniques, the OpenKruise Agents project, ecosystem integration, and production case studies with performance metrics.

AI agentsCheckpointContainer Sandbox
0 likes · 16 min read
Building Containerized Sandboxes for Multi‑Agent AI: Architecture, Key Technologies, and Real‑World Practices
Golang Shines
Golang Shines
Aug 8, 2026 · Operations

Diagnosing Server Connectivity Issues with Ping, Telnet, Curl, and Traceroute

This guide explains how to break down the vague symptom “network unreachable” into layered checks—interface status, routing, ARP, DNS, TCP, TLS, and HTTP—using the four classic tools ping, telnet, curl, and traceroute, and provides concrete commands, analysis steps, and evidence‑gathering scripts for Linux servers and Kubernetes pods.

KubernetesLinuxcurl
0 likes · 35 min read
Diagnosing Server Connectivity Issues with Ping, Telnet, Curl, and Traceroute
Cloud Architecture
Cloud Architecture
Aug 6, 2026 · Big Data

Exporting 10 Billion Elasticsearch Records: From Simple Script to Enterprise Offline Platform

The article analyses why exporting billions of Elasticsearch documents requires a full‑stack platform rather than a one‑off script, detailing the pitfalls of naive pagination, the benefits of PIT + search_after + slicing, and a complete architecture with Kafka, Redis, MySQL, Kubernetes and observability for reliable, scalable offline data export.

Data ExportElasticsearchJava
0 likes · 40 min read
Exporting 10 Billion Elasticsearch Records: From Simple Script to Enterprise Offline Platform
MaGe Linux Operations
MaGe Linux Operations
Aug 6, 2026 · Databases

How to Determine the Right Database Connection Pool Size: Practical Guidelines and Benchmarks

This article walks through a systematic approach to sizing PostgreSQL connection pools for Java applications using HikariCP and Spring Boot, covering capacity budgeting, workload‑driven calculations, monitoring metrics, slow‑SQL analysis, leak detection, Kubernetes deployment considerations, and safe rollout practices.

HikariCPKubernetesPostgreSQL
0 likes · 27 min read
How to Determine the Right Database Connection Pool Size: Practical Guidelines and Benchmarks
Cloud Native Technology Community
Cloud Native Technology Community
Aug 6, 2026 · Cloud Native

5 Production Challenges for Running AI Workloads on Kubernetes: From GPU Scheduling to Observability

Running AI workloads on Kubernetes introduces five production‑grade challenges—complex GPU and accelerator management, workload‑aware scheduling, inference autoscaling beyond CPU metrics, multi‑layer observability, and Day 2 governance—requiring platform teams to extend their capabilities beyond traditional container operations.

AI workloadsCloud NativeDay 2 operations
0 likes · 10 min read
5 Production Challenges for Running AI Workloads on Kubernetes: From GPU Scheduling to Observability
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Aug 6, 2026 · Cloud Native

How ACK Pro Provisioned Control Plane Eliminates Kubernetes Control‑Plane Bottlenecks for Large‑Scale Clusters

ACK Pro introduces a provisioned control‑plane mode that replaces reactive scaling with preset performance tiers, guaranteeing deterministic capacity for thousands of nodes and tens of thousands of Pods, and a real‑world AI training case shows reduced pod‑startup latency, eliminated HTTP 429 errors, and about 30% faster training cycles.

ACK ProAI workloadsKubernetes
0 likes · 10 min read
How ACK Pro Provisioned Control Plane Eliminates Kubernetes Control‑Plane Bottlenecks for Large‑Scale Clusters
Cloud Architecture
Cloud Architecture
Aug 5, 2026 · Backend Development

From Commit Standards to K8s Deployment: A Practical Git Engineering Guide for Backend Teams

The article explains how backend teams can achieve safe, traceable, and continuously controllable production releases for large microservice systems by building a Git‑centric engineering pipeline that covers commit conventions, branch strategies, automated versioning, immutable artifacts, GitOps configuration, and progressive canary rollouts on Kubernetes.

CI/CDGitGitOps
0 likes · 36 min read
From Commit Standards to K8s Deployment: A Practical Git Engineering Guide for Backend Teams
Raymond Ops
Raymond Ops
Aug 5, 2026 · Cloud Native

How to Diagnose Kubernetes Node NotReady Issues: A Complete Step‑by‑Step Troubleshooting Guide

This guide walks Kubernetes operators through a systematic, step‑by‑step process for diagnosing nodes stuck in the NotReady state, covering kubelet status reporting, common failure reasons, detailed command‑line checks, root‑cause analysis, remediation steps, verification, and long‑term preventive measures.

KubernetesNotReadycertificate
0 likes · 42 min read
How to Diagnose Kubernetes Node NotReady Issues: A Complete Step‑by‑Step Troubleshooting Guide
AI Open-Source Efficiency Guide
AI Open-Source Efficiency Guide
Aug 5, 2026 · Artificial Intelligence

Orchard: Microsoft’s Open‑Source Agent Framework Hits 0.28 s Latency with 1,000 Sandboxes

Orchard is Microsoft’s open‑source, Kubernetes‑native agent modeling platform that isolates execution in lightweight sandboxes, separates control‑plane operations, supports arbitrary base images and multiple built‑in harnesses, and—according to official benchmarks—delivers an average command latency of 0.28 seconds when running 1,000 concurrent sandboxes.

AI agentsKubernetesOrchard
0 likes · 17 min read
Orchard: Microsoft’s Open‑Source Agent Framework Hits 0.28 s Latency with 1,000 Sandboxes
21CTO
21CTO
Aug 4, 2026 · Backend Development

How Zalando Achieved 1M RPS with an In‑Process Client‑Side Load Balancer

Zalando’s engineering team redesigned its high‑throughput product‑read API by moving 100‑fold internal fan‑out routing into an in‑process client‑side load balancer, cutting tail latency, reducing infrastructure costs by over 75%, and improving observability while keeping the external Skipper edge router unchanged.

KubernetesMicroservicesZalando
0 likes · 7 min read
How Zalando Achieved 1M RPS with an In‑Process Client‑Side Load Balancer
Cloud Architecture
Cloud Architecture
Aug 4, 2026 · Cloud Native

From 10 to 1000 Deployments a Day: A Practical Guide to High‑Frequency Kubernetes CI/CD Architecture

This article analyses why traditional CI/CD pipelines become a bottleneck when services grow to hundreds, outlines a production‑grade, four‑plane architecture for Kubernetes that delivers declarative, auditable, concurrent, rollback‑able, gray‑scale, and extensible high‑frequency deployments, and provides concrete examples, code snippets, and a step‑by‑step rollout plan.

ArgoCDCI/CDCanary
0 likes · 37 min read
From 10 to 1000 Deployments a Day: A Practical Guide to High‑Frequency Kubernetes CI/CD Architecture
Ray's Galactic Tech
Ray's Galactic Tech
Aug 4, 2026 · Backend Development

Spring Boot & Netty MQTT Platform: Multi‑Protocol and Modular Design

This guide walks through building a high‑performance, scalable MQTT access gateway for IoT using Spring Boot for service orchestration and Netty for connection handling, covering protocol fundamentals, modular architecture, multi‑protocol adaptation, session management, high‑concurrency optimizations, clustering, observability, and deployment best practices.

IoTKubernetesMQTT
0 likes · 40 min read
Spring Boot & Netty MQTT Platform: Multi‑Protocol and Modular Design
Random Bulletin
Random Bulletin
Aug 4, 2026 · Cloud Native

Stateless Design at Ten‑Million QPS: Moving State Out of Compute for Elastic Scaling

The article analyzes how local state—such as in‑memory sessions, caches, files, and long‑lived connections—breaks elastic scaling, outlines four concrete paths to externalize that state (Redis, JWT, object storage, and gateway separation), and explains the resulting scalability benefits, trade‑offs, and performance considerations.

KubernetesMicroservicesRedis
0 likes · 20 min read
Stateless Design at Ten‑Million QPS: Moving State Out of Compute for Elastic Scaling
MaGe Linux Operations
MaGe Linux Operations
Aug 4, 2026 · Operations

How to Diagnose a 100% CPU Spike in 3 Minutes: From Symptom to Root Cause

When a production service shows 100% CPU usage, this guide walks you through a rapid three‑minute workflow—identifying the affected scope, distinguishing user, system, iowait, steal and softirq metrics, and using Linux, systemd, Docker/Kubernetes and language‑specific tools to pinpoint the offending process, thread, or system call before taking corrective actions such as throttling, scaling, or rolling back.

CPU troubleshootingDockerKubernetes
0 likes · 16 min read
How to Diagnose a 100% CPU Spike in 3 Minutes: From Symptom to Root Cause
Random Bulletin
Random Bulletin
Aug 3, 2026 · Operations

Cost Optimization at Ten‑Million QPS: Turning Ignored Expenses into Core Design

At ten‑million QPS scale, the article explains why cost shifts from a hidden after‑the‑fact bill to a primary design goal, detailing how to make costs observable, improve utilization, right‑size resources, leverage Spot and reserved instances, apply architectural savings, and embed FinOps culture while preserving SLA.

Cloud ComputingFinOpsKubernetes
0 likes · 22 min read
Cost Optimization at Ten‑Million QPS: Turning Ignored Expenses into Core Design
Ray's Galactic Tech
Ray's Galactic Tech
Aug 3, 2026 · Information Security

Are You Implementing Field-Level Encryption Correctly? Best Practices Explained

This article examines common pitfalls and misconceptions in field‑level encryption, explains why AES‑GCM and proper AAD are essential, outlines threat modeling, key hierarchy, blind indexing for searchable data, migration strategies, key rotation, performance considerations, and secure integration with MyBatis, Kubernetes, and Vault.

AES-GCMBlind IndexingField-Level Encryption
0 likes · 49 min read
Are You Implementing Field-Level Encryption Correctly? Best Practices Explained
Golang Shines
Golang Shines
Aug 3, 2026 · Cloud Native

How I Built a Production‑Ready HA Kubernetes Cluster in Minutes

When my manager suddenly demanded a production‑grade, highly available Kubernetes cluster integrated with a private Harbor registry, I followed a comprehensive step‑by‑step guide to finish the entire setup within a few hours, and now share the 83‑page manual for anyone to replicate.

Cloud NativeCluster DeploymentHarbor
0 likes · 3 min read
How I Built a Production‑Ready HA Kubernetes Cluster in Minutes
Java Architect Handbook
Java Architect Handbook
Aug 3, 2026 · Databases

Redis Officially Launches RedisInsight: A Stunning GUI with Powerful Features

RedisInsight is a visual GUI for Redis that uniquely supports Redis Cluster, offers SSL/TLS connections, memory analysis and an integrated CLI; the article walks through downloading the package, configuring environment variables, starting the service on Linux, deploying it on Kubernetes with a YAML manifest, and using the UI to monitor and operate Redis instances.

GUIInstallationKubernetes
0 likes · 8 min read
Redis Officially Launches RedisInsight: A Stunning GUI with Powerful Features
Random Bulletin
Random Bulletin
Aug 2, 2026 · Cloud Native

From Zero to Ten‑Million QPS: How to Reserve Resources for Guaranteed Capacity

The article explains how a noisy‑neighbor batch job can cripple a payment service at ten‑million‑QPS scale, then details five practical reservation techniques—quotas, reserved instances, isolation, priority preemption, and elastic prediction—while weighing their trade‑offs and showing how overcommit and offline mixing recover idle capacity for both high guarantee and high utilization.

Kubernetesovercommitpriority preemption
0 likes · 23 min read
From Zero to Ten‑Million QPS: How to Reserve Resources for Guaranteed Capacity
Golang Shines
Golang Shines
Aug 2, 2026 · Cloud Native

GPU Scheduling, Isolation, and Resource Allocation in Kubernetes Clusters

This guide explains why a GPU‑enabled node may show devices with nvidia‑smi yet keep Pods pending, walks through the complete node‑to‑container GPU path, and provides step‑by‑step procedures for device discovery, Device Plugin configuration, isolation models (full‑card, time‑slicing, MIG), scheduling constraints, quota management, multi‑GPU training, troubleshooting pending Pods, and monitoring with DCGM metrics.

Device PluginKubernetesMIG
0 likes · 33 min read
GPU Scheduling, Isolation, and Resource Allocation in Kubernetes Clusters
Architect Chen
Architect Chen
Aug 2, 2026 · Cloud Native

All Essential kubectl Commands for 2026: A Complete Guide

This article provides a concise, step‑by‑step reference of the most frequently used kubectl commands—including get, describe, logs, exec, apply, port‑forward, rollout, scale, and delete—showing exact syntax and typical use cases for managing Kubernetes resources.

Cloud NativeContainer ManagementDevOps
0 likes · 4 min read
All Essential kubectl Commands for 2026: A Complete Guide
MaGe Linux Operations
MaGe Linux Operations
Aug 2, 2026 · Cloud Native

K8s Multi‑Tenant Isolation: Practical Hierarchical Namespaces with Namespace + HNC

This guide explains how to achieve robust multi‑tenant isolation in a shared Kubernetes cluster by combining Namespace with the Hierarchical Namespace Controller (HNC), covering isolation dimensions, permission and network policies, resource quotas, hierarchy design, verification steps, and high‑risk operation safeguards.

HNCKubernetesNetworkPolicy
0 likes · 29 min read
K8s Multi‑Tenant Isolation: Practical Hierarchical Namespaces with Namespace + HNC
Random Bulletin
Random Bulletin
Jul 31, 2026 · Cloud Native

Scaling at Ten‑Million QPS: From Manual to Automatic Autoscaling

The article analyzes why manual capacity adjustments break down at ten‑million‑QPS scale, then walks through metric‑driven autoscaling, anti‑flapping algorithms, headroom planning, predictive scaling, stateful service challenges, and multi‑dimensional strategies to achieve a cost‑stable dynamic balance.

AutoscalingCloud NativeKubernetes
0 likes · 20 min read
Scaling at Ten‑Million QPS: From Manual to Automatic Autoscaling
MaGe Linux Operations
MaGe Linux Operations
Jul 31, 2026 · Cloud Native

Advanced Kubernetes Scheduling: Pod Affinity, Anti‑Affinity, and Topology Spread Constraints

This article explains how node affinity, pod affinity/anti‑affinity, and topologySpreadConstraints differ, shows how to diagnose pending Pods, provides best‑practice YAML examples, and integrates these rules with PDBs, PriorityClasses, metrics, and rollback procedures for reliable, highly available workloads.

KubernetesNodeAffinityPDB
0 likes · 27 min read
Advanced Kubernetes Scheduling: Pod Affinity, Anti‑Affinity, and Topology Spread Constraints
MaGe Linux Operations
MaGe Linux Operations
Jul 31, 2026 · Cloud Native

Choosing an Ingress Controller: Production Comparison of NGINX, Traefik, and APISIX

This article presents a production‑grade comparison of three Kubernetes Ingress controllers—NGINX, Traefik, and APISIX—by defining a four‑layer evaluation framework, detailing pre‑deployment checks, configuration examples, testing scripts, performance metrics, and rollout/rollback procedures to help teams select the most suitable solution.

APISIXCloud NativeKubernetes
0 likes · 24 min read
Choosing an Ingress Controller: Production Comparison of NGINX, Traefik, and APISIX
Xiaolin Talks Programming
Xiaolin Talks Programming
Jul 30, 2026 · Cloud Native

Spring Boot 3.x + Spring Cloud Kubernetes: Native Service Discovery, Config Injection & Graceful Shutdown

This production-hardened guide shows how to replace Eureka/Nacos with Kubernetes-native Service/EndpointSlice for service discovery, use ConfigMap/Secret for configuration with dynamic refresh, and implement zero-downtime deployments via preStop hooks and Spring Boot graceful shutdown — complete with RBAC hardening, fault-drill baselines, and a 7-point production checklist.

ConfigMapGraceful ShutdownKubernetes
0 likes · 16 min read
Spring Boot 3.x + Spring Cloud Kubernetes: Native Service Discovery, Config Injection & Graceful Shutdown
DevOps Operations Practice
DevOps Operations Practice
Jul 30, 2026 · Operations

Essential Velero Guide for Kubernetes Disaster Recovery

This article walks through using Velero to back up, restore, and migrate Kubernetes clusters, covering MinIO installation, Velero client and server setup, storage volume creation, backup location configuration, and execution of backup, restore, and scheduled backup commands.

Cloud NativeKubernetesOperations
0 likes · 11 min read
Essential Velero Guide for Kubernetes Disaster Recovery