Tagged articles

Kubernetes

4439 articles · Page 2 of 45
DevOps Operations Practice
DevOps Operations Practice
Jul 30, 2026 · Operations

Essential Velero Guide for Kubernetes Disaster Recovery

This article walks through using Velero to back up, restore, and migrate Kubernetes clusters, covering MinIO installation, Velero client and server setup, storage volume creation, backup location configuration, and execution of backup, restore, and scheduled backup commands.

Cloud NativeKubernetesOperations
0 likes · 11 min read
Essential Velero Guide for Kubernetes Disaster Recovery
PaperAgent
PaperAgent
Jul 29, 2026 · Artificial Intelligence

How to Build Harness‑Native Agents Using OpenForge RL

OpenForge RL introduces a lightweight proxy and Kubernetes‑based orchestrator to decouple training from inference, enabling the training of 30B‑scale and 8B agents within any harness, while providing an automatic five‑stage task synthesis pipeline and demonstrating state‑of‑the‑art results across Claw, GUI, and Browser benchmarks.

AgentHarnessKubernetes
0 likes · 13 min read
How to Build Harness‑Native Agents Using OpenForge RL
Full-Stack DevOps & Kubernetes
Full-Stack DevOps & Kubernetes
Jul 29, 2026 · Operations

How AI-Integrated EFK Lets Machines Handle Log Screening and Fault Diagnosis

The article examines how integrating large language models with the EFK logging stack transforms traditional, manual log inspection into an AI-driven process that automatically filters logs, identifies anomalies, performs root‑cause analysis, and generates structured fault reports, dramatically improving operational efficiency and reducing mean‑time‑to‑resolution.

AI OpsAIOpsEFK
0 likes · 10 min read
How AI-Integrated EFK Lets Machines Handle Log Screening and Fault Diagnosis
Golang Shines
Golang Shines
Jul 28, 2026 · Cloud Native

Building a Fully Automated GitOps Delivery Pipeline with Argo CD on Kubernetes

This guide walks through implementing a GitOps workflow for Kubernetes using Argo CD, covering repository structure, Helm chart management, permission boundaries, Argo CD installation, application manifests, CI integration, troubleshooting OutOfSync states, and handover procedures, all with concrete commands and examples.

Argo CDCI/CDGit
0 likes · 30 min read
Building a Fully Automated GitOps Delivery Pipeline with Argo CD on Kubernetes
Xiaolin Talks Programming
Xiaolin Talks Programming
Jul 28, 2026 · Cloud Native

Production-Grade Spring Boot Containerization: Layered Images, JVM Tuning & OOM Debugging

A comprehensive guide to running Spring Boot reliably on Kubernetes covering layered Docker image builds, JVM container-aware memory and CPU configuration, OOMKilled root-cause analysis using Native Memory Tracking and heap dumps, Prometheus monitoring integration, and Cloud Native Buildpacks for automated CI/CD pipelines.

BuildpacksCI/CDDocker
0 likes · 18 min read
Production-Grade Spring Boot Containerization: Layered Images, JVM Tuning & OOM Debugging
Cloud Architecture
Cloud Architecture
Jul 27, 2026 · Cloud Native

Kubernetes Pod Storage Panorama: Full Chain from Volume, PVC, PV to CSI

The article walks through the entire Kubernetes storage lifecycle, explaining how Volumes, PersistentVolumeClaims, PersistentVolumes, StorageClasses and CSI interact, and shows real‑world production scenarios, common pitfalls, and practical guidance for designing reliable, scalable storage solutions for stateful workloads.

CSIDevOpsKubernetes
0 likes · 37 min read
Kubernetes Pod Storage Panorama: Full Chain from Volume, PVC, PV to CSI
dbaplus Community
dbaplus Community
Jul 26, 2026 · Cloud Native

Will AI Replace Kubernetes? Co‑Founder Brendan Burns on Its Rise and End

Brendan Burns recounts how he convinced Google to back Kubernetes, built the MVP in five days, navigated open‑source governance, tackled technical challenges like Etcd and declarative design, expanded the platform for AI workloads, and reflects on why even successful software like Kubernetes inevitably faces obsolescence.

AI workloadsCloud NativeKubernetes
0 likes · 31 min read
Will AI Replace Kubernetes? Co‑Founder Brendan Burns on Its Rise and End
Black & White Path
Black & White Path
Jul 26, 2026 · Information Security

CVE‑2026 PoC Collection: All 12 Exploit Codes in One Repository

A GitHub repository named cve-2026-poc-collection, maintained by researcher XZ1r0, aggregates twelve high‑severity CVE‑2026 proof‑of‑concept exploits across web, Linux, Windows and other categories, detailing their impact, implementation languages, and offering a search script to help security professionals quickly locate and analyze the code.

CVE-2026KubernetesPoC
0 likes · 9 min read
CVE‑2026 PoC Collection: All 12 Exploit Codes in One Repository
DataFunSummit
DataFunSummit
Jul 25, 2026 · Cloud Native

Evolution of Agent Infrastructure: Engineering Insights from Tencent Cloud Agent Runtime

The article analyzes how agents transition from demo to production, revealing that beyond model capabilities, stability, elasticity, security, and governance become critical, and explains the engineering challenges and solutions—including session management, state persistence, scheduling mismatches, sandbox isolation, and open‑source strategies—that underpin Tencent Cloud's Agent Runtime.

Agent RuntimeCloud NativeKubernetes
0 likes · 26 min read
Evolution of Agent Infrastructure: Engineering Insights from Tencent Cloud Agent Runtime
CodeSmart Hoops
CodeSmart Hoops
Jul 25, 2026 · Backend Development

Interview Self‑Test: Spring Boot Actuator Health Checks – Quick Review & Must‑Know Answers

This article provides a comprehensive interview self‑test covering Spring Boot Actuator’s built‑in endpoints, health‑check JSON format, status aggregation logic, custom HealthIndicator implementation, show‑details configuration, @Endpoint operation annotations, liveness vs. readiness probes, handling of DOWN status, disabling auto‑configured indicators, security risks of exposing all endpoints, HealthIndicatorRegistry naming rules, additional Actuator endpoints, CompositeHealth construction, and custom HealthStatusHttpMapper mapping.

ActuatorJavaKubernetes
0 likes · 22 min read
Interview Self‑Test: Spring Boot Actuator Health Checks – Quick Review & Must‑Know Answers
Ops Development Stories
Ops Development Stories
Jul 25, 2026 · Cloud Native

Practical Guide to Pyrra: The Kubernetes‑Native SLO Monitoring Tool

This comprehensive guide explains how Pyrra extends Sloth by providing a full SLO platform for Kubernetes, covering its architecture, four SLI types, rule generation, Web UI features, alert configuration, deployment options, Grafana integration, advanced usage, common pitfalls, and a detailed comparison to help you choose the right tool for reliable service monitoring.

Cloud NativeKubernetesPrometheus
0 likes · 24 min read
Practical Guide to Pyrra: The Kubernetes‑Native SLO Monitoring Tool
Raymond Ops
Raymond Ops
Jul 24, 2026 · Cloud Native

How to Properly View Container Logs Without Using Tail -f Inside Pods

This article explains the correct ways to view container logs in Kubernetes, covering the underlying storage mechanism, kubectl log commands, log rotation, distributed log collection architectures like EFK and Loki, best‑practice recommendations, and detailed troubleshooting steps for common log‑related issues.

EFKKubernetesLoki
0 likes · 28 min read
How to Properly View Container Logs Without Using Tail -f Inside Pods
Geek Labs
Geek Labs
Jul 24, 2026 · Artificial Intelligence

How Google’s New Open‑Source Projects Make AI Agents Production‑Ready

Google Cloud recently open‑sourced two Go projects—Scion, which isolates and coordinates multiple AI agents, and AX, a distributed runtime that enables a single long‑running agent to resume after failures—detailing their architectures, usage steps, real‑world use cases, limitations, and the broader strategy of turning agents from experimental toys into reliable production workers.

AI AgentsAXGo
0 likes · 12 min read
How Google’s New Open‑Source Projects Make AI Agents Production‑Ready
Ray's Galactic Tech
Ray's Galactic Tech
Jul 22, 2026 · Backend Development

From Chaos to Control: A Complete Spring Boot Logging Guide with 8 Production Scenarios

The article analyzes why logging in modern Spring Boot microservices shifts from a simple debug tool to a critical data pipeline, compares Logback and Log4j2 async setups, reveals hidden costs and risks, and walks through eight concrete production scenarios to build a controllable, high‑performance logging architecture.

KubernetesLogbackMicroservices
0 likes · 25 min read
From Chaos to Control: A Complete Spring Boot Logging Guide with 8 Production Scenarios
dbaplus Community
dbaplus Community
Jul 21, 2026 · Databases

Tired of Hand‑Crafted Backup Scripts? Meet Databasus – One Open‑Source Platform for All Major Databases

Databasus is an open‑source backup management platform that unifies PostgreSQL, MySQL, MariaDB and MongoDB backups with a web UI, offering logical, physical and incremental backups, automated scheduling, AES‑256‑GCM encryption, multi‑cloud storage, RBAC collaboration, and optional agents for secure, out‑bound connections.

AES-256-GCMDatabasusDocker
0 likes · 14 min read
Tired of Hand‑Crafted Backup Scripts? Meet Databasus – One Open‑Source Platform for All Major Databases
Cloud Architecture
Cloud Architecture
Jul 21, 2026 · Cloud Native

Kubernetes Troubleshooting in Practice: 20 Survival Rules from Real Incidents

This article presents a hands‑on guide to diagnosing Kubernetes production failures, distilling a real e‑commerce outage into 20 actionable rules that cover nodes, control plane, networking, scheduling, storage and observability, and provides a step‑by‑step diagnostic workflow with concrete commands and examples.

ConfigMapHPAKubernetes
0 likes · 29 min read
Kubernetes Troubleshooting in Practice: 20 Survival Rules from Real Incidents
Ray's Galactic Tech
Ray's Galactic Tech
Jul 21, 2026 · Backend Development

14 Painful Spring Boot Cache Pitfalls and How to Build a High‑Availability Architecture

This article walks through 14 real‑world failure scenarios of Spring Boot distributed caching, explains why high cache hit rates are misleading, and provides concrete analysis, code samples, and step‑by‑step recommendations for designing a resilient cache layer that isolates faults, handles hot keys, and ensures data consistency across Redis, local caches, and databases.

Cache invalidationKubernetesPerformance
0 likes · 37 min read
14 Painful Spring Boot Cache Pitfalls and How to Build a High‑Availability Architecture
MaGe Linux Operations
MaGe Linux Operations
Jul 21, 2026 · Cloud Native

How to Use Kubernetes Node Affinity to Schedule Large Models on Specific GPU Nodes

This guide explains how to schedule large‑model inference pods onto GPU nodes that meet exact hardware requirements—such as A100 80 GB cards, specific node pools, and zones—by converting those needs into Kubernetes node‑affinity, taint, and topology constraints, verifying the deployment, monitoring its health, and safely rolling out or rolling back changes.

GPU SchedulingKubernetesNode Affinity
0 likes · 23 min read
How to Use Kubernetes Node Affinity to Schedule Large Models on Specific GPU Nodes
MaGe Linux Operations
MaGe Linux Operations
Jul 21, 2026 · Cloud Native

Auto‑Scaling LLM Inference with Kubernetes HPA Based on Request Queue Depth

The article explains how to replace CPU‑only autoscaling for large‑model inference services with a Kubernetes HPA that scales pods according to a custom queue‑depth metric exported to Prometheus, covering metric definition, deployment configuration, Prometheus‑Adapter setup, HPA creation, capacity calculation, validation, troubleshooting, and rollback procedures.

AutoscalingHPAKubernetes
0 likes · 21 min read
Auto‑Scaling LLM Inference with Kubernetes HPA Based on Request Queue Depth
TechVision Expert Circle
TechVision Expert Circle
Jul 21, 2026 · Cloud Native

How to Build an Elastic Auto‑Scaling Cloud‑Native Application

After a 15‑fold traffic surge forced manual scaling of an e‑commerce platform, the team rebuilt the system with true elastic scaling—horizontal, vertical, and architectural—using Kubernetes, Envoy, KEDA, predictive autoscaling, and a comprehensive observability stack, achieving fully automated scaling from 12 to 80 pods in under 90 seconds and cutting peak resource costs by 60%.

Cloud NativeElastic ScalingKEDA
0 likes · 13 min read
How to Build an Elastic Auto‑Scaling Cloud‑Native Application
Cloud Architecture
Cloud Architecture
Jul 20, 2026 · Cloud Native

Kubernetes Authentication Time Bomb: The Evolution and Production Practices of ServiceAccount Tokens

The article explains how many teams mistakenly think they are using Kubernetes authentication while actually mounting long‑lived Bearer tokens, outlines the risks of legacy ServiceAccount tokens, describes the new projected token mechanism, and provides step‑by‑step guidance for secure production deployment and migration.

KubernetesProjectedVolumeServiceAccount
0 likes · 21 min read
Kubernetes Authentication Time Bomb: The Evolution and Production Practices of ServiceAccount Tokens
MaGe Linux Operations
MaGe Linux Operations
Jul 19, 2026 · Operations

Hands‑On nvidia‑smi Guide: Diagnosing GPU Utilization and Memory Usage Anomalies

This article provides a step‑by‑step, Linux‑focused workflow for recording driver and GPU versions, interpreting utilization versus memory metrics, locating memory‑consuming processes, handling container and Kubernetes mappings, checking temperature, power, ECC, MIG, driver health, OOM conditions, and setting up reliable monitoring and alert thresholds for data‑center GPUs.

CUDAGPU monitoringKubernetes
0 likes · 28 min read
Hands‑On nvidia‑smi Guide: Diagnosing GPU Utilization and Memory Usage Anomalies
Golang Shines
Golang Shines
Jul 18, 2026 · Cloud Native

Mastering GPU Scheduling, Isolation, and Resource Allocation in Kubernetes Clusters

This guide walks through the complete GPU resource path from node to container, explains how Kubernetes discovers and registers GPUs via the NVIDIA Device Plugin, and provides step‑by‑step procedures for environment inventory, pod specifications, scheduling constraints, isolation models, quota management, monitoring, troubleshooting, and safe upgrades.

Device PluginKubernetesMIG
0 likes · 36 min read
Mastering GPU Scheduling, Isolation, and Resource Allocation in Kubernetes Clusters
Ops Community
Ops Community
Jul 18, 2026 · Operations

Using tcpdump to Diagnose Client‑Server Communication Failures

This guide shows how to use tcpdump on Linux to verify whether a client actually sent packets, whether they reached the server, how the server responded, and where a TCP connection was closed, by defining the problem, selecting interfaces, applying narrow BPF filters, capturing key handshake packets, handling TLS, HTTP, UDP, DNS, ICMP, container and Kubernetes environments, and preserving evidence with proper file management.

KubernetesLinuxTCP
0 likes · 19 min read
Using tcpdump to Diagnose Client‑Server Communication Failures
Linyb Geek Road
Linyb Geek Road
Jul 18, 2026 · Cloud Native

How to Accelerate Spring Boot Startup on Kubernetes Using CRaC

This guide walks through enabling CRaC in a Spring Boot application, building a CRaC‑compatible Docker image, creating a checkpoint job in Kubernetes, restoring the snapshot at pod start, and comparing the startup speed with GraalVM native compilation.

Azul ZuluCRaCDocker
0 likes · 11 min read
How to Accelerate Spring Boot Startup on Kubernetes Using CRaC
Cloud Architecture
Cloud Architecture
Jul 17, 2026 · Cloud Native

Stop Hand‑Crafting ClusterRoles: Build a Production‑Grade Kubernetes RBAC Governance System with rbac‑manager

This article explains why manually managing ClusterRoles leads to governance chaos in Kubernetes, introduces rbac‑manager as a declarative controller that centralises binding creation, recycling and auditing, and provides a step‑by‑step guide with real‑world examples to build a scalable, production‑ready RBAC management workflow.

Access ControlCloud NativeKubernetes
0 likes · 23 min read
Stop Hand‑Crafting ClusterRoles: Build a Production‑Grade Kubernetes RBAC Governance System with rbac‑manager
Golang Shines
Golang Shines
Jul 17, 2026 · Cloud Native

Building a Scalable Go Service Mesh from Scratch: Core Cloud‑Native Practices

This article walks through why Go is ideal for cloud‑native development and demonstrates step‑by‑step how to build a scalable service mesh, covering static compilation, HTTP services, Go modules, Gin/Gorilla APIs, configuration, logging, health checks, service registration, load balancing, sidecar proxies, traffic interception, circuit breaking, rate limiting, retries, and distributed tracing with OpenTelemetry.

Cloud NativeGoKubernetes
0 likes · 16 min read
Building a Scalable Go Service Mesh from Scratch: Core Cloud‑Native Practices
Ops Community
Ops Community
Jul 17, 2026 · Cloud Native

Monitoring GPU Metrics with DCGM Exporter and Prometheus

This guide explains how to continuously monitor NVIDIA GPU utilization, memory, temperature, power and error metrics using DCGM Exporter, covering driver verification, Docker and Compose deployment, Prometheus scraping, Kubernetes DaemonSet setup, custom collectors, PromQL queries, alert rules and troubleshooting procedures.

AlertingDCGM ExporterDocker
0 likes · 28 min read
Monitoring GPU Metrics with DCGM Exporter and Prometheus
Cloud Architecture
Cloud Architecture
Jul 16, 2026 · Cloud Native

Mastering the Kubernetes Control Plane: From Informer Source Code to a Production‑Ready Dynamic Gateway Operator

The article explains why naïve operators that only watch a few resources quickly fail under load, then dives into the true purpose of the Informer pipeline, demonstrates how to design a four‑layer state model for a dynamic gateway, and provides production‑grade patterns for reconciliation, status handling, governance, and when an Operator is truly needed.

CRDControl PlaneDynamic Gateway
0 likes · 25 min read
Mastering the Kubernetes Control Plane: From Informer Source Code to a Production‑Ready Dynamic Gateway Operator
Ray's Galactic Tech
Ray's Galactic Tech
Jul 16, 2026 · Artificial Intelligence

K8s, Kafka, Nacos Agent Platform to Prevent Token Bankruptcy and Skill Avalanches

The article details how a production‑grade Agent platform built on Kubernetes, Kafka, and Nacos addresses token budget overruns, uncontrolled skill execution, and RAG hallucinations by introducing a four‑layer runtime architecture, token pre‑allocation, explicit state management, dynamic governance policies, and robust skill specifications.

KafkaKubernetesLLM Agents
0 likes · 32 min read
K8s, Kafka, Nacos Agent Platform to Prevent Token Bankruptcy and Skill Avalanches
Ops Community
Ops Community
Jul 16, 2026 · Cloud Native

How to Use Kubernetes PVC for Persistent Pod Storage

This guide explains why persistent storage is essential for Kubernetes Pods, details the responsibilities of PVC, PV, StorageClass and CSI, and provides step‑by‑step commands, checks, and best‑practice procedures for creating, troubleshooting, expanding, migrating, and safely deleting PVCs in production environments.

CSIDataMigrationKubernetes
0 likes · 39 min read
How to Use Kubernetes PVC for Persistent Pod Storage
TechVision Expert Circle
TechVision Expert Circle
Jul 15, 2026 · Cloud Native

Designing a Live‑Streaming Platform for 1.2 Million Concurrent Viewers

To support 1.2 million simultaneous viewers, the article details a three‑layer push‑stream‑transcode‑distribution architecture, SRT/WHIP protocols, AV1 GPU‑accelerated transcoding, multi‑CDN edge delivery, a scalable WebSocket message system, Kubernetes‑based auto‑scaling, and extensive performance tuning and disaster‑recovery strategies.

AV1CDNKubernetes
0 likes · 13 min read
Designing a Live‑Streaming Platform for 1.2 Million Concurrent Viewers
MaGe Linux Operations
MaGe Linux Operations
Jul 15, 2026 · Cloud Native

How to Schedule, Isolate, and Allocate GPUs in a Kubernetes Cluster

Even when GPU nodes show up with nvidia‑smi, Pods can stay pending, see all devices, or suffer memory spikes; this guide walks through the full GPU resource chain in Kubernetes, from PCIe detection and driver loading to Device Plugin registration, node labeling, affinity, taints, isolation levels, MIG, time‑slicing, quotas, monitoring, and safe upgrade procedures.

Cloud NativeDevice PluginKubernetes
0 likes · 34 min read
How to Schedule, Isolate, and Allocate GPUs in a Kubernetes Cluster
Cloud Architecture
Cloud Architecture
Jul 14, 2026 · Operations

From Avalanche to Self‑Healing: Why Nginx 502 Spikes During High‑Traffic Sales and How to Fix It

During large‑scale promotions a sudden flood of Nginx 502 errors signals upstream interaction failures across proxy, kernel, application and orchestration layers, and the article explains the exact conditions, root causes, traffic amplification, and a systematic self‑healing approach to diagnose and eliminate them.

502KubernetesNginx
0 likes · 27 min read
From Avalanche to Self‑Healing: Why Nginx 502 Spikes During High‑Traffic Sales and How to Fix It
Ray's Galactic Tech
Ray's Galactic Tech
Jul 14, 2026 · Cloud Native

Spring Boot + Netty MQTT Gateway for Million Connections & Millisecond Push

To support millions of persistent MQTT connections with sub‑millisecond latency, the article walks through a Spring Boot + Netty cloud‑native gateway design that separates connection, event, state and governance planes, details async authentication, back‑pressure handling, command state machines, and loss‑less Kubernetes roll‑outs.

Cloud NativeKafkaKubernetes
0 likes · 38 min read
Spring Boot + Netty MQTT Gateway for Million Connections & Millisecond Push
Raymond Ops
Raymond Ops
Jul 14, 2026 · Cloud Native

Kubernetes Networking: From CNI Basics to Troubleshooting

This article explains Kubernetes' three‑principle network model, compares the leading CNI plugins (Flannel, Calico, Cilium), details pod communication paths, Service and Ingress mechanisms, DNS and NetworkPolicy implementations, and provides step‑by‑step troubleshooting cases with performance data and concrete configuration examples.

CNICalicoCilium
0 likes · 34 min read
Kubernetes Networking: From CNI Basics to Troubleshooting
MaGe Linux Operations
MaGe Linux Operations
Jul 14, 2026 · Databases

Common MySQL Connection Errors and Step‑by‑Step Troubleshooting Guide

MySQL connection failures are among the most frequent issues for developers and operators; this article systematically walks through typical error messages, explains how to collect relevant information, runs layered command checks, analyzes evidence, identifies root causes such as socket problems, bind‑address limits, host whitelist mismatches, authentication failures, connection‑limit exhaustion, and packet timeouts, and provides concrete fix and verification procedures for on‑premise, Docker, and Kubernetes deployments.

DockerKubernetesLinux
0 likes · 25 min read
Common MySQL Connection Errors and Step‑by‑Step Troubleshooting Guide
Golang Shines
Golang Shines
Jul 14, 2026 · Cloud Native

How to Build a Kubernetes Cluster from Scratch: Step‑by‑Step Guide

This article walks you through planning, hardware preparation, system initialization, Docker and kubeadm installation, certificate generation, etcd deployment, master and node component configuration, CNI networking, TLS bootstrapping, and final verification to create a fully functional Kubernetes cluster from the ground up.

CNIDockerKubernetes
0 likes · 26 min read
How to Build a Kubernetes Cluster from Scratch: Step‑by‑Step Guide
Golang Shines
Golang Shines
Jul 14, 2026 · Operations

Why Does OOM Occur Even When Server Memory Looks Sufficient?

The article explains that out‑of‑memory (OOM) events can happen despite apparent free memory because OOM can be triggered by cgroup limits, NUMA constraints, kernel allocation failures, or systemd‑oomd policies, and it provides a step‑by‑step diagnostic method covering logs, metrics, and Kubernetes specifics.

KubernetesLinuxMemory
0 likes · 27 min read
Why Does OOM Occur Even When Server Memory Looks Sufficient?
360 Zhihui Cloud Developer
360 Zhihui Cloud Developer
Jul 14, 2026 · Cloud Native

Using Pod Overhead and Kata Containers to Isolate Kernel Memory and Stop Container Slab Leaks

The article explains how intensive file‑system reads in a Kubernetes pod cause kernel slab memory to balloon, why standard cgroup limits cannot contain the leak, and demonstrates step‑by‑step how configuring Pod Overhead with Kata containers creates a separate sandbox cgroup that isolates kernel memory, preventing host‑level OOM.

Kata ContainersKubernetesPod Overhead
0 likes · 12 min read
Using Pod Overhead and Kata Containers to Isolate Kernel Memory and Stop Container Slab Leaks
java1234
java1234
Jul 14, 2026 · Backend Development

Quarkus: Java Framework Up to 5× Faster Than Spring Boot, Uses Half the Memory

Quarkus, Red Hat's cloud‑native Java framework, achieves dramatically faster startup (0.3‑1 s vs 2‑5 s) and roughly half the memory usage (150‑200 MB vs 300‑500 MB) by moving many optimizations to compile time, and offers a Spring‑like developer experience with extensions, reactive support, and hot‑reload tooling.

Dev ModeJavaKubernetes
0 likes · 9 min read
Quarkus: Java Framework Up to 5× Faster Than Spring Boot, Uses Half the Memory
Cloud Architecture
Cloud Architecture
Jul 13, 2026 · Cloud Native

Ultimate Guide to Choosing the Right kube-proxy Mode for Production Kubernetes

This comprehensive guide explains how kube-proxy drives Service traffic in Kubernetes, compares userspace, iptables, IPVS, nftables and eBPF modes, and provides a four‑dimensional decision framework, migration steps, monitoring practices, and real‑world examples to help operators select the optimal mode for their clusters.

IPVSKubernetesNetworking
0 likes · 25 min read
Ultimate Guide to Choosing the Right kube-proxy Mode for Production Kubernetes
Ray's Galactic Tech
Ray's Galactic Tech
Jul 13, 2026 · Artificial Intelligence

When AI Agents Meet Cloud‑Native: Practical Multi‑Agent Orchestration for High‑Concurrency Scenarios

The article explains why naïve multi‑agent demos fail in production, defines the core concepts of Task, Step, Agent Role and Event, proposes a four‑plane cloud‑native architecture, shows concrete Go and Python code, and provides detailed guidance on state machines, reliability, observability, security and budget governance for building scalable, production‑grade AI agent systems.

AI AgentsCloud NativeKubernetes
0 likes · 36 min read
When AI Agents Meet Cloud‑Native: Practical Multi‑Agent Orchestration for High‑Concurrency Scenarios
Raymond Ops
Raymond Ops
Jul 13, 2026 · Operations

Scaling Prometheus to Thousands of Nodes with Thanos: Architecture, Storage, and HA Practices

The article analyzes the storage, query performance, high‑availability, and data‑loss challenges of running Prometheus on a 1,000‑node Kubernetes cluster and demonstrates how a Thanos‑based architecture—Sidecar, Query, Store Gateway, Compactor, Receiver, and object‑storage back‑ends—can be designed, tuned, and operated to achieve horizontal scalability, efficient down‑sampling, and reliable fault recovery.

KubernetesObject StoragePrometheus
0 likes · 35 min read
Scaling Prometheus to Thousands of Nodes with Thanos: Architecture, Storage, and HA Practices
Golang Shines
Golang Shines
Jul 12, 2026 · Operations

10 Essential Linux Ops Tools Every Engineer Should Master

This article introduces ten widely used Linux operations tools—Shell scripts, Git, Ansible, Prometheus, Grafana, Docker, Kubernetes, Nginx, ELK Stack, and Zabbix—detailing their functions, typical scenarios, advantages, concrete usage examples, and links to learning resources for each.

DockerGitGrafana
0 likes · 9 min read
10 Essential Linux Ops Tools Every Engineer Should Master
MaGe Linux Operations
MaGe Linux Operations
Jul 12, 2026 · Operations

Why Does OOM Occur Even When Server Memory Looks Sufficient?

Even when monitoring shows free memory, Linux can still kill processes due to various OOM paths such as cgroup limits, NUMA allocation failures, kernel high-order allocation issues, or systemd‑oomd, and this guide walks through a reproducible investigation and remediation process.

KubernetesLinuxOOM
0 likes · 29 min read
Why Does OOM Occur Even When Server Memory Looks Sufficient?
Raymond Ops
Raymond Ops
Jul 11, 2026 · Cloud Native

Kubernetes HPA & VPA Auto-Scaling: Elastic Strategies for Traffic Spikes

An in‑depth comparison of Kubernetes Horizontal and Vertical Pod Autoscalers—including algorithms, configurations, performance benchmarks, mixed‑mode trade‑offs, custom‑metric integrations, and real‑world case studies—demonstrates how to choose and tune HPA, VPA, and KEDA for rapid traffic spikes while avoiding conflicts.

AutoscalingCloud NativeHPA
0 likes · 47 min read
Kubernetes HPA & VPA Auto-Scaling: Elastic Strategies for Traffic Spikes
MaGe Linux Operations
MaGe Linux Operations
Jul 11, 2026 · Operations

Step‑by‑Step Guide to Diagnose 100 % CPU on a Linux Server

When a Linux server’s CPU spikes to 100 %, this article walks through a systematic investigation—from defining what “CPU 100 %” really means, gathering timestamps and metrics, using tools like top, mpstat, vmstat, pidstat, sar, perf, and strace, to tracing processes, threads, containers, and Kubernetes, building an evidence chain, applying low‑risk fixes, and verifying the resolution.

CPUKubernetesLinux
0 likes · 24 min read
Step‑by‑Step Guide to Diagnose 100 % CPU on a Linux Server
Ops Community
Ops Community
Jul 10, 2026 · Operations

How to Diagnose Network Packet Loss: Practical Steps from ping to tcpdump

This guide explains how to systematically investigate network packet loss on Linux by collecting timestamps, routing, interface and kernel statistics, using ping, tracepath, curl, netstat, ss, nstat, ip, ethtool, nftables/iptables, and tcpdump on both ends, then narrowing the failure scope, fixing MTU, conntrack, soft‑interrupt or firewall issues, and validating the fix with metrics and roll‑back procedures.

KubernetesLinuxconntrack
0 likes · 18 min read
How to Diagnose Network Packet Loss: Practical Steps from ping to tcpdump
Ctrip Technology
Ctrip Technology
Jul 10, 2026 · Cloud Native

How Ctrip Scaled Karmada to Handle 200 GB Memory Peaks and Hundreds of Thousands of Pods

This article details Ctrip's migration from Kubefed to Karmada for multi‑cluster governance, describing the architectural evolution, production deployment for high‑availability and smooth cross‑cluster pod migration, and the performance and memory optimizations required to support over 200 GB of control‑plane memory usage and tens of thousands of Pods.

Control Plane OptimizationFederationKarmada
0 likes · 18 min read
How Ctrip Scaled Karmada to Handle 200 GB Memory Peaks and Hundreds of Thousands of Pods
Raymond Ops
Raymond Ops
Jul 9, 2026 · Operations

Practical Guide to Troubleshooting and Resolving DNS Issues

This comprehensive guide explains how DNS works, categorises common resolution failures, and provides step‑by‑step procedures, command‑line examples and configuration snippets for diagnosing and fixing DNS problems in Linux, Kubernetes and cloud environments.

DNSKubernetesTroubleshooting
0 likes · 52 min read
Practical Guide to Troubleshooting and Resolving DNS Issues
TechVision Expert Circle
TechVision Expert Circle
Jul 8, 2026 · Backend Development

Designing a 200M‑Request Recommendation System with 50ms P99 Latency

The team rebuilt a recommendation platform for a 80‑million‑DAU content service, scaling daily requests from 30 M to 200 M, cutting P99 latency from 800 ms to under 50 ms by introducing a four‑layer architecture, multi‑path recall (vector, real‑time, graph), transformer‑based ranking, multi‑level caching, predictive autoscaling, and comprehensive observability.

Feature StoreKubernetesVector Search
0 likes · 13 min read
Designing a 200M‑Request Recommendation System with 50ms P99 Latency
Cloud Architecture
Cloud Architecture
Jul 7, 2026 · Databases

MySQL User & Permission Management: From Grant Statements to Production-Grade Security Architecture

This comprehensive guide explains why MySQL permission mistakes happen, walks through the authentication and authorization process, shows how to design multi‑layered user models, role hierarchies, declarative GitOps workflows, Kubernetes integration, and production‑ready automation for secure, auditable, and scalable database access.

GitOpsKubernetesMySQL
0 likes · 42 min read
MySQL User & Permission Management: From Grant Statements to Production-Grade Security Architecture
Golang Shines
Golang Shines
Jul 7, 2026 · Operations

Mastering Linux Server Time Synchronization with NTP and Chrony: Best Practices

This guide explains why time synchronization is a critical yet often overlooked part of Linux operations, outlines common failure scenarios, and provides a step‑by‑step methodology for configuring, verifying, and troubleshooting NTP/chrony across physical servers, virtual machines, containers, and Kubernetes clusters.

KubernetesLinuxNTP
0 likes · 42 min read
Mastering Linux Server Time Synchronization with NTP and Chrony: Best Practices
dbaplus Community
dbaplus Community
Jul 6, 2026 · Cloud Native

Why Companies Are Switching from Kubernetes to K3s: Simplicity, Stability, and Low Overhead

The article explains how K3s, a lightweight, production‑grade Kubernetes distribution, reduces component count, memory usage, and installation complexity, making it ideal for small‑to‑medium enterprises, edge, IoT, and AI projects, while still offering API compatibility and optional high‑availability, and compares its trade‑offs with full‑size Kubernetes.

InstallationKubernetesedge computing
0 likes · 7 min read
Why Companies Are Switching from Kubernetes to K3s: Simplicity, Stability, and Low Overhead
Cloud Architecture
Cloud Architecture
Jul 6, 2026 · Backend Development

Spring Boot Template Engine Mix: Designing a Multi‑Engine Coexistence Architecture for Production

The article explains why running multiple template engines in a Spring Boot application becomes an operational challenge, outlines a four‑layer architecture and routing strategies, provides concrete code for a custom ViewResolver, configuration, observability, deployment, testing and migration steps, and shows how to govern the process safely in production.

KubernetesSpring BootThymeleaf
0 likes · 39 min read
Spring Boot Template Engine Mix: Designing a Multi‑Engine Coexistence Architecture for Production
Java Tech Enthusiast
Java Tech Enthusiast
Jul 6, 2026 · Cloud Native

Why Use Service Registry & Discovery When Nginx Already Handles Load Balancing?

The article analyzes Nginx's static upstream load balancing limitations—manual configuration, passive health checks, and inability to handle elastic scaling—and explains how service registries provide real‑time instance awareness, client‑side load balancing, metadata‑driven routing, and seamless scaling for microservices.

KubernetesMicroservicesNginx
0 likes · 9 min read
Why Use Service Registry & Discovery When Nginx Already Handles Load Balancing?
Java Architect Handbook
Java Architect Handbook
Jul 6, 2026 · Cloud Native

Why We Dropped Nacos for Apollo: A Hands‑On Guide to Configuration Management

This article explains why the team replaced Nacos with Ctrip's open‑source Apollo configuration center, outlines Apollo's core concepts, features, and architecture, and provides step‑by‑step instructions for creating projects, testing dynamic updates, exploring environments, clusters, namespaces, and deploying a SpringBoot application on Kubernetes.

Configuration CenterJavaKubernetes
0 likes · 28 min read
Why We Dropped Nacos for Apollo: A Hands‑On Guide to Configuration Management
Full-Stack DevOps & Kubernetes
Full-Stack DevOps & Kubernetes
Jul 6, 2026 · Cloud Native

Taming Massive Alert Noise: A Hands‑On Guide to AI‑Driven Dynamic Thresholds for Prometheus

This article presents a practical solution that uses Facebook Prophet time‑series AI to automatically calibrate dynamic alert thresholds in Prometheus, reducing over‑80% of false alarms in Kubernetes environments by learning business cycles and updating rules hourly without manual intervention.

AIOpsDynamic ThresholdFacebook Prophet
0 likes · 10 min read
Taming Massive Alert Noise: A Hands‑On Guide to AI‑Driven Dynamic Thresholds for Prometheus
Ops Community
Ops Community
Jul 5, 2026 · Operations

20 Common Ops Newbie Pitfalls – Which Ones Have You Hit?

This guide catalogs the 20 most frequent mistakes made by new operations engineers, explains why they happen, and provides step‑by‑step safe alternatives, risk warnings, and recovery procedures so readers can avoid costly outages and build reliable habits.

DevOpsKubernetesLinux
0 likes · 29 min read
20 Common Ops Newbie Pitfalls – Which Ones Have You Hit?
ThinkingAgent
ThinkingAgent
Jul 4, 2026 · Cloud Native

Building the AI Infra Foundation: L0 Resource Layer for GPU Scheduling and Cloud‑Native Architecture

The article presents a detailed, step‑by‑step analysis of the L0 resource layer that underpins AI infrastructure, covering GPU scheduling, multi‑tier storage, low‑latency networking, core architectural components, key technologies such as MIG, Volcano, Kueue and RDMA, practical implementation patterns, quantitative acceptance criteria, and common pitfalls with best‑practice mitigations.

AI InfrastructureCloud NativeGPU Scheduling
0 likes · 26 min read
Building the AI Infra Foundation: L0 Resource Layer for GPU Scheduling and Cloud‑Native Architecture
MaGe Linux Operations
MaGe Linux Operations
Jul 4, 2026 · Operations

20 Common Ops Rookie Mistakes and How to Avoid Them

This guide lists the twenty most frequent pitfalls that new operations engineers encounter, explains why they happen, and provides step‑by‑step safe practices, code examples, risk classifications and a verification checklist to help prevent costly outages and data loss.

DevOpsKubernetesLinux
0 likes · 28 min read
20 Common Ops Rookie Mistakes and How to Avoid Them
Cloud Architecture
Cloud Architecture
Jul 2, 2026 · Databases

MySQL Containerization vs Host Installation: Production‑Grade Selection Framework

The article explains that the real challenge is not merely running MySQL but placing it in the right resource model, and it provides a four‑dimensional decision framework—performance ceiling, stability floor, automation level, and organizational maturity—to guide when to use host‑installed MySQL, single‑node containers, or full Kubernetes deployment, illustrated with concrete resource analyses, architecture diagrams, configuration examples, pitfalls, checklists, and an evolution roadmap.

KubernetesMySQLPerformance
0 likes · 35 min read
MySQL Containerization vs Host Installation: Production‑Grade Selection Framework
Architect Chen
Architect Chen
Jul 2, 2026 · Cloud Native

A Complete Visual Guide to Kubernetes Architecture

This article provides a comprehensive, step‑by‑step overview of Kubernetes architecture, detailing the control plane components (API Server, etcd, Scheduler, Controller Manager) and worker node components (kubelet, kube‑proxy, container runtimes), illustrated with diagrams and command‑line examples that show how requests flow, state is stored, pods are scheduled, and failures are handled.

Control PlaneController ManagerK8s Architecture
0 likes · 4 min read
A Complete Visual Guide to Kubernetes Architecture
Raymond Ops
Raymond Ops
Jul 1, 2026 · Operations

Memory Leak Postmortem: Combining free, smem, pmap, and perf for Effective Diagnosis

When a thumbnail service experienced sudden latency spikes and OOM kills shortly after a new release, the author walks through a systematic investigation using free, smem, pmap, and perf to distinguish true memory leaks from page‑cache or shared‑page artifacts, pinpoint the native decoder buffer issue, and outline remediation steps.

KubernetesLinuxMemory Leak
0 likes · 29 min read
Memory Leak Postmortem: Combining free, smem, pmap, and perf for Effective Diagnosis
Golang Shines
Golang Shines
Jul 1, 2026 · Operations

10 Essential Ops Tools That Can Cut Your Overtime by 80%

This article introduces ten Linux operations tools—Shell scripts, Git, Ansible, Prometheus, Grafana, Docker, Kubernetes, Nginx, ELK Stack, and Zabbix—detailing their functions, typical use cases, advantages, and concrete examples to help engineers streamline daily tasks and dramatically reduce overtime.

DockerGitGrafana
0 likes · 9 min read
10 Essential Ops Tools That Can Cut Your Overtime by 80%
Xiaolin Talks Programming
Xiaolin Talks Programming
Jun 30, 2026 · Databases

Zero-Downtime Database Evolution: Flyway + Spring Boot Expand/Migrate/Contract Pattern

This article details a production-verified zero-downtime database migration strategy using Spring Boot and Flyway, covering why DDL causes outages, Flyway configuration best practices, the expand-migrate-contract pattern for smooth schema changes, script writing and rollback techniques, Kubernetes deployment coordination, monitoring with circuit breakers, and a post-mortem checklist.

DDLFlywayKubernetes
0 likes · 18 min read
Zero-Downtime Database Evolution: Flyway + Spring Boot Expand/Migrate/Contract Pattern
dbaplus Community
dbaplus Community
Jun 29, 2026 · Cloud Computing

Why More Companies Are Dropping VMware for Proxmox

Since 2024, a growing number of enterprises—especially small‑to‑medium businesses and some large firms—are re‑evaluating the cost‑driven VMware licensing model and migrating to the open‑source Proxmox VE platform, which bundles KVM, LXC, Ceph, backup and clustering into a free, easy‑to‑manage solution that fits modern AI and Kubernetes workloads.

Cloud NativeKubernetesProxmox
0 likes · 6 min read
Why More Companies Are Dropping VMware for Proxmox
Raymond Ops
Raymond Ops
Jun 28, 2026 · Operations

Why Large‑Model Services Keep Running Out of GPU Memory: An Ops View from KV Cache to Concurrency

The article explains why large‑model inference services frequently hit GPU memory limits, breaks down static vs. dynamic memory consumption, shows how KV‑Cache, request length, and concurrency amplify usage, and provides a step‑by‑step troubleshooting and mitigation workflow for production environments.

GPU memoryInference OptimizationKV Cache
0 likes · 26 min read
Why Large‑Model Services Keep Running Out of GPU Memory: An Ops View from KV Cache to Concurrency
Architect's Guide
Architect's Guide
Jun 28, 2026 · Cloud Native

Kubernetes Networking Explained with 16 Detailed Diagrams

This article provides a comprehensive, diagram‑driven analysis of Kubernetes networking, covering underlay and overlay models, the role of VLAN, OSPF, BGP, and various CNI plugins such as Flannel host‑gw, Calico BGP, IPVLAN/MACVLAN, Multus, and Danm, as well as tunnel technologies like VxLAN and IPIP.

CNICalicoFlannel
0 likes · 13 min read
Kubernetes Networking Explained with 16 Detailed Diagrams
Raymond Ops
Raymond Ops
Jun 27, 2026 · Operations

Hands‑On DNS Ops: Deploy BIND and CoreDNS with Full Troubleshooting Guide

This comprehensive guide walks you through DNS fundamentals, compares BIND, CoreDNS, PowerDNS and Unbound, provides step‑by‑step deployment scripts for BIND 9.20 and CoreDNS 1.12, explains DNSSEC configuration, caching optimizations, security hardening, high‑availability designs, monitoring, backup and recovery procedures, and advanced troubleshooting techniques.

BINDCoreDNSDNS
0 likes · 43 min read
Hands‑On DNS Ops: Deploy BIND and CoreDNS with Full Troubleshooting Guide
Golang Shines
Golang Shines
Jun 26, 2026 · Cloud Native

Why Every Ops Role Now Demands Kubernetes Skills (And a 100‑Question K8s Interview Guide)

After being laid off after five years in operations, the author realized that all job listings now require Docker and Kubernetes expertise, so they compiled a comprehensive "100 K8s Interview Questions" guide covering core concepts, architecture, resource management, networking, storage, security, troubleshooting, and ecosystem tools.

Cloud NativeContainer OrchestrationDevOps
0 likes · 7 min read
Why Every Ops Role Now Demands Kubernetes Skills (And a 100‑Question K8s Interview Guide)
Alibaba Cloud Infrastructure
Alibaba Cloud Infrastructure
Jun 26, 2026 · Cloud Computing

How Kimi’s AI Agent Scales on Alibaba Cloud – Architecture, Elastic Sandbox, and Cost Optimisation

The article analyses how Kimi’s AI Agent workloads are deployed on Alibaba Cloud using ACK and the ACS Agent Sandbox, detailing the challenges of massive concurrency, rapid sandbox start‑up, state continuity, cost‑effective scaling, and the security and scheduling mechanisms that enable production‑grade performance.

AI AgentAlibaba CloudCost Optimisation
0 likes · 19 min read
How Kimi’s AI Agent Scales on Alibaba Cloud – Architecture, Elastic Sandbox, and Cost Optimisation
Subtle Storm
Subtle Storm
Jun 25, 2026 · Backend Development

Why Microservices Matter: Core Architecture and Key Technologies Explained

The article analyzes how monolithic applications struggle with scalability, deployment, and team coordination, then breaks down microservice architecture—including API gateways, service communication, registration, resilience patterns, tracing, and container orchestration—while weighing its benefits against added complexity and team size considerations.

ArchitectureCircuit BreakerDocker
0 likes · 7 min read
Why Microservices Matter: Core Architecture and Key Technologies Explained
Architect Chen
Architect Chen
Jun 25, 2026 · Cloud Native

Four Key Ways to Deploy Microservices: From Bare Metal to Kubernetes

The article compares four microservice deployment approaches—physical servers, virtual machines, containerization with Docker, and Kubernetes clusters—detailing their implementation, advantages, drawbacks, and ideal scenarios, helping teams choose the most suitable strategy based on resource isolation, scalability, operational complexity, and team expertise.

Cloud NativeKubernetesMicroservices
0 likes · 6 min read
Four Key Ways to Deploy Microservices: From Bare Metal to Kubernetes
Ops Development & AI Practice
Ops Development & AI Practice
Jun 24, 2026 · Information Security

Ending Hard‑Coded Rules: OPA Policy‑as‑Code for Unified SecOps Guardrails

The article explains how enterprises can replace fragmented, hard‑coded security checks in Terraform, CI/CD pipelines, Kubernetes admission webhooks, and API gateways with a unified, declarative policy engine—Open Policy Agent—using Rego to decouple decision and enforcement, enabling fast, auditable SecOps guardrails across the entire software lifecycle.

CI/CDKubernetesOPA
0 likes · 12 min read
Ending Hard‑Coded Rules: OPA Policy‑as‑Code for Unified SecOps Guardrails
Cloud Architecture
Cloud Architecture
Jun 22, 2026 · Backend Development

Dubbo vs Spring Cloud: Deep Dive for Billion‑Scale Microservice Architecture

The article examines how to choose between Dubbo and Spring Cloud for high‑traffic microservice systems, analyzing communication models, thread and connection handling, governance capabilities, real‑world e‑commerce scenarios, and provides practical guidance on combining HTTP gateways, RPC, and asynchronous messaging for scalable, resilient architectures.

DubboKubernetesMicroservices
0 likes · 32 min read
Dubbo vs Spring Cloud: Deep Dive for Billion‑Scale Microservice Architecture
Raymond Ops
Raymond Ops
Jun 22, 2026 · Artificial Intelligence

Elastic Deployment and GPU Scheduling for Large‑Model Inference with vLLM on Kubernetes

This article presents a detailed, step‑by‑step analysis of deploying the high‑performance vLLM inference engine on Kubernetes, covering GPU memory management, tensor parallelism, quantization choices, continuous batching, and automated scaling with HPA/KEDA to achieve low latency and high throughput for large language models.

DockerGPU SchedulingKubernetes
0 likes · 49 min read
Elastic Deployment and GPU Scheduling for Large‑Model Inference with vLLM on Kubernetes
Ops Development Stories
Ops Development Stories
Jun 22, 2026 · Cloud Native

Design and Implementation of a Multi‑Cluster Arthas‑Based Online Diagnosis Platform

This article details the architecture, security mechanisms, and implementation of a unified Arthas online diagnosis platform that enables SSH‑free, audited access to Java applications across dozens of isolated Kubernetes clusters, covering control‑plane design, WebSocket tunneling, credential management, RBAC, and front‑end integration with Vue and xterm.js.

ArthasCloud NativeGo
0 likes · 23 min read
Design and Implementation of a Multi‑Cluster Arthas‑Based Online Diagnosis Platform
Raymond Ops
Raymond Ops
Jun 21, 2026 · Cloud Native

Stop Pods From “Running Wild”: A Practical Guide to Kubernetes Scheduling Strategies

This guide explains why default Kubernetes scheduling often falls short in production, introduces nodeSelector, nodeAffinity, podAffinity/anti‑affinity, taints/tolerations, topologySpreadConstraints and PriorityClass, and provides step‑by‑step configuration examples, real‑world use cases, best‑practice recommendations, troubleshooting tips, and monitoring alerts to ensure reliable pod placement.

KubernetesNodeAffinityPod Scheduling
0 likes · 36 min read
Stop Pods From “Running Wild”: A Practical Guide to Kubernetes Scheduling Strategies
Raymond Ops
Raymond Ops
Jun 20, 2026 · Operations

Eliminate Monitoring Blind Spots: Hands‑On Enterprise‑Grade Prometheus + Grafana Deployment

This comprehensive guide walks you through the end‑to‑end setup of a production‑grade Prometheus and Grafana monitoring stack, covering architecture choices, installation steps, configuration details, high‑availability designs, performance tuning, security hardening, troubleshooting, backup strategies, and best‑practice recommendations.

AlertingGrafanaKubernetes
0 likes · 49 min read
Eliminate Monitoring Blind Spots: Hands‑On Enterprise‑Grade Prometheus + Grafana Deployment
DataFunTalk
DataFunTalk
Jun 19, 2026 · Artificial Intelligence

How NVIDIA Dynamo Boosts Multi‑Node Distributed Inference MFU for Agentic AI

The article explains how NVIDIA Dynamo tackles the production bottlenecks of Agentic AI by using KV‑Cache‑aware routing, a three‑stage multimodal inference architecture, and intelligent cache scheduling on Kubernetes to improve multi‑node throughput (MFU) while maintaining latency SLAs.

Agentic AIKV CacheKubernetes
0 likes · 3 min read
How NVIDIA Dynamo Boosts Multi‑Node Distributed Inference MFU for Agentic AI
Programmer XiaoFu
Programmer XiaoFu
Jun 18, 2026 · Cloud Native

Why Use Service Registration When Nginx Already Handles Load Balancing?

The article explains that Nginx’s static upstream configuration and passive health checks cannot keep up with dynamic microservice environments, while a service registry provides real‑time instance awareness, automatic failure detection, and metadata‑driven routing, making both tools complementary rather than interchangeable.

EurekaKubernetesNacos
0 likes · 9 min read
Why Use Service Registration When Nginx Already Handles Load Balancing?
Architecture & Thinking
Architecture & Thinking
Jun 18, 2026 · Backend Development

How to Scale a Flash‑Sale System from Zero to 1 Million QPS: A Step‑by‑Step Architecture Guide

This article dissects the evolution of a flash‑sale system from a simple monolithic controller to a cloud‑native, micro‑service architecture that can handle over one million requests per second, detailing traffic‑shaping, multi‑level caching, async processing, and inventory‑consistency techniques.

Distributed ArchitectureKubernetesRedis
0 likes · 18 min read
How to Scale a Flash‑Sale System from Zero to 1 Million QPS: A Step‑by‑Step Architecture Guide