Cloud Native Not Dead: Kubernetes, CNCF, Cilium Redesign Foundations for AI Agents
Tony Bai analyzes how Kubernetes, CNCF, Cilium, and OpenTelemetry are fundamentally redesigning cloud-native primitives—scheduling, resource allocation, sandbox runtimes, gateways, and observability—to support stateful, bursty AI agents, proving cloud-native isn't obsolete but evolving its foundational interfaces.
Refuting the "Cloud Native Is Obsolete" Narrative
The article opens with three common arguments for cloud-native obsolescence: agents are stateful and bursty unlike microservices; Kubernetes operational burden (clusters, registries, CRDs) deters developers; and new runtimes like Agent Substrate bypass Kubernetes' control plane. However, the author counters with concrete evidence from major project updates over the past 21 months: Kubernetes v1.37 (67 enhancements), Cilium 1.20 (2,660+ commits), OpenTelemetry CNCF graduation (May 2026), Kubernetes AI Conformance platforms growing from 18 to 31 with agent workload verification, and the deliberate retirement of Ingress NGINX in favor of Gateway API. Notably, even Substrate still uses Kubernetes as its infrastructure provisioning layer and aligns with upstream Pod certificate capabilities.
Layered View of Cloud-Native Foundations for the Agent Era
The author compresses the ecosystem into a layered diagram, each layer explored in subsequent sections.
Kubernetes: Scheduling and Resource Layer Rewritten
Five Version Snapshots (v1.33–v1.37)
v1.33 (Apr 2025, Octarine): User namespaces default, In-Place Resize Beta, Image Volumes Beta.
v1.34 (Aug 2025, Of Wind & Will): DRA GA, Pod-level resources Beta, PSI metrics Beta.
v1.35 (Dec 2025, Timbernetes): In-Place Resize GA, first workload-aware scheduling.
v1.36 (Apr 2026, Haru): User namespaces GA, topology-aware scheduling and workload-aware preemption (Alpha), manifest-based admission control.
v1.37 (Aug 2026, Garhwal): Gang Scheduling Beta, DRA Extended Resource GA, HPA Scale to Zero Beta, Metrics API stable.
Scheduling Object Shifts from Pod to Workload
AI training and distributed inference require a group of Pods to run together or not at all. Traditional per-Pod scheduling causes deadlocks where resources are held but the group cannot form. Kubernetes introduces Workload-Aware Scheduling (WAS) with CompositePodGroup (v1.37) organizing PodGroups into a tree supporting multi-level topology constraints, gang scheduling, and preemption strategies. Official callouts include JobSet and LeaderWorkerSet. Caution: API versions changed from v1alpha1 to v1alpha2 (v1.36) to v1alpha3 (v1.37) with a breaking disruptionMode field; KEP-4671 targets v1.38 for stable, but production use should be validated in pre-production first.
DRA: Standard Device Abstraction for the GPU Era
Dynamic Resource Allocation (DRA) reached GA in v1.34; v1.37 adds critical pieces:
Extended Resource GA: Traditional requests like example.com/gpu can be fulfilled by DRA drivers without Device Plugin coexistence, enabling gradual migration.
Device Taints and Tolerations GA: Overheated or maintenance-bound GPUs can be tainted with NoSchedule (block new scheduling) and NoExecute (evict running Pods).
Standard NUMA Node Attributes: Enables cross-vendor device comparison on the same NUMA node.
PodGroup Shared ResourceClaim (Beta): Avoids each Pod in a group creating its own Claim.
Elasticity and Isolation for Long-Lived, Bursty Agents
In-Place Pod Resize GA in v1.35; v1.37 adds preemption to make room for in-place expansion (Alpha).
HPA Scale to Zero enters Beta in v1.37.
User Namespaces GA in v1.36: container root no longer equals host root.
Memory QoS Beta and default-on in v1.37.
Manifest-based admission control enables immutable platform security baselines.
Pod Certificates and Cluster Trust Bundles establish workload identity.
Clean Break: Ingress NGINX Retirement
Announced Nov 2025, joint steering committee statement Jan 2026, maintenance stopped Mar 2026. Official Ingress2Gateway 1.0 released for migration. Community moved the traffic entrance standard squarely onto Gateway API, which continues to advance: v1.5 (Apr 2026) multiple features to Stable, v1.6 (Aug 2026) TCPRoute and UDPRoute to Standard. All subsequent AI gateway capabilities build on this foundation.
Google AX and Agent Substrate: The Most Radical Redesign
Why Pods Are Insufficient
Kubernetes blog "Running Agents on Kubernetes with Agent Sandbox" describes "AI v2 eating AI v1": from short-lived stateless calls to continuously running multi-agent collaboration. Agents have three traits that break traditional primitives:
Singleton and stateful: need persistent identity and a secure "scratchpad" for executing untrusted LLM-generated code.
Mostly idle, occasional bursts: need suspend and fast resume.
Startup latency sensitive: a new Pod adds ~1 second overhead, breaking interaction continuity when an agent wakes.
Combining StatefulSet, Headless Service, and PVC works at small scale but becomes an operational nightmare at scale.
Path One: Agent Sandbox (Improving Within the Pod Abstraction)
Developed under SIG Apps, core is the Sandbox CRD:
Strong isolation: native support for gVisor and Kata Containers.
Lifecycle management: scale to zero when idle, resume seamlessly.
Stable identity: stable hostname and network identity for inter-agent discovery.
Extension layer adds three CRDs: SandboxTemplate (template), SandboxClaim (on-demand request), SandboxWarmPool (pre-warmed pool to reduce cold start). Google announced GKE Agent Sandbox technical preview at KubeCon NA 2025.
Path Two: Agent Substrate and AX (Moving Control Plane Out of Hot Path)
GKE documentation argues standard Kubernetes binds each agent to a separate Pod, limited by scheduling throughput and Pod startup latency; Kubernetes also lacks Pod hibernation, so millions of idle agents would exhaust Pod limits and control-plane memory. Substrate multiplexes many "Actors" (agents) onto few ready "Workers" (Pods). When idle, the agent's memory and local files are snapshotted; on demand, restored into an available sandbox in under a second.
AX is a Kubernetes-style declarative orchestration layer atop Substrate (Apache 2.0), defining four primitives:
Task: execution lifecycle, sandbox resource constraints.
Workspace: pre-execution environment setup: mount Git repos, configure MCP servers, install Skill packages.
Gateway: egress network policy: hostname/port allowlists and credential injection.
Model: unified entry for LLM provider parameters, runtime config, and secrets.
CLI resembles kubectl:
ax apply -f task.yaml # register manifest
ax watch # real-time task phase and status changes
ax ssh <task> # debug inside sandbox
ax suspend <task> # manual suspend
ax resume <task> # resumeA telling detail: Substrate repo includes a Pod certificate signing controller described as a "patch-style" implementation whose capability will eventually be upstreamed into Kubernetes—aligning with v1.37's Pod Certificates. The most radical agent runtime is converging toward upstream standards, not forking.
Addressing the Strongest Criticisms
Hacker News debate: infrastructure engineers praise solving "idle agents burning money"; developers criticize "ergonomics" claims vs. Kubernetes cluster/registry/CRD operational burden.
Co-creator Jaana Dogan emphasizes AX is job orchestration, not an agent framework .
AX Go docs marked active early development; Substrate README states not an officially supported Google product; deployment docs target GKE Standard clusters.
These criticisms show Kubernetes isn't the only answer for agents, but most teams' compute, network, security, observability, and identity stacks are built on cloud-native; bypassing it is costly. The accurate statement: the agent-era foundation isn't torn down but disassembled and re-layered.
Traffic Layer: From Header to Payload
Traditional gateways route by headers; AI requests carry key info in the body: model, priority, prompt safety.
Jun 2025: Gateway API Inference Extension released, using Envoy ext-proc to upgrade any Gateway API-compliant gateway into an inference gateway. Project states GA.
Mar 2026: SIG Network forms AI Gateway Working Group to add AI-oriented declarative capabilities to Gateway API. Two core proposals:
Payload Processing: declaratively process full request/response bodies for prompt injection protection, content filtering, semantic routing, caching, RAG integration; requires deterministic processor ordering, configurable failure modes, explicit MCP request handling.
Egress: outbound routing to external model providers (OpenAI, Vertex AI, Bedrock).
Agent Gateway: agentgateway (Rust, created by Solo.io, now Linux Foundation) covers MCP, A2A, and LLM routing; split from kgateway as independent project May 2026. kagent (CNCF Sandbox May 2025) uses CRDs to declare agents, model configs, and MCP tools.
Inference Layer: llm-d entered CNCF Sandbox Mar 24, 2026, initiated by Red Hat, Google Cloud, IBM Research, CoreWeave, NVIDIA, etc. Built on vLLM and Inference Extension, implements Prefill/Decode separation and prefix-cache-aware routing. Google blog discloses this routing raised Vertex AI prefix cache hit rate from 35% to 70%.
CNCF: Binding the Ecosystem with "Contracts"
Kubernetes AI Conformance (launched Nov 2025) requires standard Kubernetes conformance plus AI-specific requirements. Mar 2026 updates signal:
Certified platforms grew from 18 to 31.
New KAR (Kubernetes AI Requirements) aligned with v1.35 primitives, mandating stable In-Place Resize and workload-aware scheduling.
Verification scope extended to Agent workloads ; standards for decoupled inference, LLM traffic routing, DRA networking in progress.
Roadmap: automated conformance testing, sovereign AI standards including sandbox and data privacy.
CNCF's "The great migration" blog (Mar 2026) frames three eras: microservices, data & training/inference, and Agent era starting 2025 . AI Tech Radar shows 41% of AI developers self-identify as cloud-native developers. Protocol governance clarifies: MCP under Linux Foundation's Agentic AI Foundation, A2A also under Linux Foundation.
Cilium: Network and Runtime Security "Cross-Cutting Layer"
Cilium 1.20 (Sep 2026) highlights two Agent-relevant tracks:
Gateway API becomes more complete traffic management layer: ExternalAuth (GEP-1494), TCPRoute/UDPRoute, tracking Gateway API v1.6.1.
Network policy standardization: upstream ClusterNetworkPolicy support with Admin and Baseline tiers; Hubble associates audit verdicts with policies.
Additional: extensible datapath (plugins, no vendor forks), ENI IPAM IPv6 Beta, ztunnel still Beta, legacy Mutual Authentication deprecated, cilium-cni binary reduced 80%. Isovalent blog Feb 2026 lists 18+ clouds defaulting to Cilium CNI (CoreWeave, OVHcloud). Tetragon Network Policy (Sep 2026) enables policy by binary match and FQDN .
OpenTelemetry: From "Three Pillars" to "Agent Call Trees"
OpenTelemetry simultaneously solidifies its base and chases agents.
Foundation
Graduated May 2026.
Declarative configuration stable.
Profiles enters public Alpha.
OBI (eBPF auto-instrumentation, donated from Grafana Beyla) advancing.
Kubernetes attributes processor v1.0.0 released Sep 2026.
Kubernetes v1.37 promotes native histograms to Beta and enables by default.
Agent Focus
Official blogs: "AI Agent Observability" (2025), "Inside the LLM Call" (2026).
Semantic conventions v1.42.0 (Jun 12, 2026) splits GenAI, provider-specific conventions, and MCP conventions into separate repo semantic-conventions-genai for faster iteration.
Conventions cover inference, embedding, retrieval, memory operations, tool execution; Agent and MCP are first-class citizens.
But overall still marked Development ; as of Aug 21, 2026 no formal release version.
Implication: Instrumenting with gen_ai.* attributes is correct, but expect attribute names to change.
End-to-End Agent Request Flow
One sentence: Gateway handles ingress, Sandbox/Substrate handles runtime, DRA and scheduler handle resources, Cilium handles boundaries, OTel handles evidence. All five occur on cloud-native turf.
Back to the Title: Who Says Cloud Native Is Obsolete?
Five Judgments
Layering, not replacement. Kubernetes core stays generic; agent-specific needs met by CRDs, sub-projects, and upper-layer systems (Sandbox, Substrate, AX).
Scheduling unit upgrading: Pod → Workload/PodGroup → Actor.
Gateway focus shifts from Header to Payload.
Standards first. AI Conformance is CNCF's "contract" to constrain the ecosystem.
Observability is the agent's "evidence chain," but standards not yet settled.
Risk Checklist
Workload-aware scheduling: API version churn v1.36→v1.37, avoid direct production use.
AI Gateway proposals lack merged CRDs; agentgateway still v1alpha1.
AX early development; Substrate currently targets GKE Standard.
OTel GenAI conventions still Development.
Ingress NGINX lesson: critical component dependent on few maintainers is inherent risk.
References
Kubernetes Blog: https://kubernetes.io/blog/
Running Agents on Kubernetes with Agent Sandbox: https://kubernetes.io/blog/2026/03/20/running-agents-on-kubernetes-with-agent-sandbox/
Kubernetes v1.37 Release: https://kubernetes.io/blog/2026/08/26/kubernetes-v1-37-release/
v1.37 Workload-Aware Scheduling: https://kubernetes.io/blog/2026/09/08/kubernetes-v1-37-advancing-workload-aware-scheduling/
v1.37 DRA Updates: https://kubernetes.io/blog/2026/09/03/kubernetes-v1-37-dra-updates/
Announcing the AI Gateway Working Group: https://kubernetes.io/blog/2026/03/09/announcing-ai-gateway-wg/
AI Gateway WG Payload Processing Proposal: https://github.com/kubernetes-sigs/wg-ai-gateway/blob/main/proposals/7-payload-processing.md
agent-sandbox repo: https://github.com/kubernetes-sigs/agent-sandbox
KEP-4671 Gang Scheduling: https://www.kubernetes.dev/resources/keps/4671/
CNCF Announcement (AI Conformance): https://www.cncf.io/announcements/2026/03/24/cncf-nearly-doubles-certified-kubernetes-ai-platforms/
CNCF Blog (The great migration): https://www.cncf.io/blog/2026/03/05/the-great-migration-why-every-ai-platform-is-converging-on-kubernetes/
CNCF Blog (Cilium 1.20): https://www.cncf.io/blog/2026/09/14/cilium-1-20-gateway-api-externalauth-tcproute-udproute-eni-ipam-for-ipv6-and-more/
Cilium 1.20.0 Release Notes: https://github.com/cilium/cilium/releases/tag/v1.20.0
OpenTelemetry Blog: https://opentelemetry.io/blog/
OpenTelemetry Graduation Announcement: https://opentelemetry.io/blog/2026/otel-graduates/
GKE: About Agent Substrate: https://docs.cloud.google.com/kubernetes-engine/ai-ml/about-agent-substrate
Google Cloud (llm-d CNCF Sandbox): https://cloud.google.com/blog/products/containers-kubernetes/llm-d-officially-a-cncf-sandbox-project
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
TonyBai
Tony Bai's tech world (tonybai.com). Not satisfied with just "knowing how", we strive for mastery. Focused on Go language internals, high-quality engineering practices, and cloud‑native architecture, exploring cutting‑edge intersections of Go and AI. Gophers who pursue technology are welcome—follow me and evolve with Go.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
