Tagged articles

PodDisruptionBudget

3 articles · Page 1 of 1
MaGe Linux Operations
MaGe Linux Operations
Jul 21, 2026 · Cloud Native

How to Use Kubernetes Node Affinity to Schedule Large Models on Specific GPU Nodes

This guide explains how to schedule large‑model inference pods onto GPU nodes that meet exact hardware requirements—such as A100 80 GB cards, specific node pools, and zones—by converting those needs into Kubernetes node‑affinity, taint, and topology constraints, verifying the deployment, monitoring its health, and safely rolling out or rolling back changes.

GPU schedulingKubernetesLarge Language Model
0 likes · 23 min read
How to Use Kubernetes Node Affinity to Schedule Large Models on Specific GPU Nodes
Raymond Ops
Raymond Ops
Jun 3, 2026 · Operations

10 Critical Kubernetes Production Failures I Caused and How to Recover

The article walks through ten real‑world Kubernetes production incidents—from an etcd disk‑full disaster to image‑pull failures—detailing symptoms, root‑cause analysis, step‑by‑step remediation commands, and preventive measures such as monitoring, quota alerts, and configuration best practices.

API ServerETCDHorizontalPodAutoscaler
0 likes · 25 min read
10 Critical Kubernetes Production Failures I Caused and How to Recover
Cloud Architecture
Cloud Architecture
May 12, 2026 · Cloud Native

Zero Downtime Isn't Accidental: Deep Dive into Kubernetes Smooth Deployments from Theory to Production

Zero‑downtime releases require coordinated control‑plane and data‑plane actions—proper RollingUpdate settings, pod lifecycle handling, service‑mesh draining, pre‑warm of dependencies, capacity safeguards, and automated monitoring/rollback—otherwise brief spikes of 502/503 errors and duplicate consumption will appear.

Argo RolloutsCanary DeploymentHPA
0 likes · 41 min read
Zero Downtime Isn't Accidental: Deep Dive into Kubernetes Smooth Deployments from Theory to Production