AI Writes Kubernetes YAML in Seconds: The Real Value of Ops Engineers
The article tests AI tools like DeepSeek for generating Kubernetes YAML, finding they handle standard templates well but fail on cluster-specific configs, security, probes, resource quotas, and complex multi-CRD scenarios. It argues ops engineers' value lies in troubleshooting, architecture decisions, incident handling, setting standards, and building platforms—not writing YAML—and advises embracing AI for drafts while deepening core expertise.
AI Writing Kubernetes YAML: How Good Is It?
A new colleague used DeepSeek to generate Deployment, Service, ConfigMap, and Ingress YAML in 30 seconds. The output included resource limits, readinessProbe, and affinity — looking plausible at first glance. This prompted the author to spend a week testing DeepSeek and ChatGPT across various Kubernetes YAML requirements.
Testing AI-Generated YAML
Conclusion: AI is already sufficient for writing standardized YAML. For example, asking for a Spring Boot Deployment with HPA produces a structurally correct manifest.
Example: Spring Boot Deployment with HPA
apiVersion: apps/v1
kind: Deployment
metadata:
name: devapp
labels:
app: devapp
spec:
replicas: 3
selector:
matchLabels:
app: devapp
template:
metadata:
labels:
app: devapp
spec:
containers:
- name: devapp
image: devapp:latest
ports:
- containerPort: 8080
resources:
requests:
cpu: 250m
memory: 512Mi
limits:
cpu: 500m
memory: 1Gi
readinessProbe:
httpGet:
path: /actuator/health
port: 8080
initialDelaySeconds: 30
periodSeconds: 10
livenessProbe:
httpGet:
path: /actuator/health
port: 8080
initialDelaySeconds: 60
periodSeconds: 30
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: devapp-hpa
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: devapp
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70It looks fine, but production-ready YAML is another matter.
Five Typical Problems with AI-Generated YAML
1. Unaware of Cluster Context
AI does not know your node configuration, CNI, or storage classes. It may suggest storageClassName: standard while your cluster uses nfs-client or local-path. It might add hostPort which can conflict with Calico/IPVS.
2. Missing Security Configurations
AI rarely adds critical security settings unless explicitly asked:
securityContext:
runAsNonRoot: true
runAsUser: 1000
readOnlyRootFilesystem: true
allowPrivilegeEscalation: falsePod-level securityContext, NetworkPolicy, and PodDisruptionBudget are often omitted.
3. Overly Templated Probe Settings
AI uses generic values like initialDelaySeconds: 30 for every service. A lightweight API gateway and a cache-heavy microservice have vastly different startup profiles; AI cannot know your application's actual behavior.
4. Arbitrary Resource Quotas
Requests like cpu: 250m, memory: 512Mi are guesses. The author saw AI assign 256Mi to a data-analysis service, causing OOMKill.
5. Inability to Handle Complex Scenarios
ArgoCD Application with blue-green deployment
Cross-namespace ServiceAccount + RBAC
Ingress with cert-manager auto TLS
StatefulSet with PodTopologySpread for multi-AZ scheduling
These multi-CRD scenarios require heavy manual rewriting.
Where Is the Real Value of Ops Engineers?
Writing YAML was never the core value. The real time sinks are:
1. Troubleshooting
When a Pod is CrashLoopBackOff, you examine logs, events, resource usage, network policies, DNS — AI cannot see real-time cluster state, topology, or recent ConfigMap changes.
2. Architecture Decisions
Deployment vs StatefulSet, need for PDB, HPA metrics (CPU vs custom), Service type (ClusterIP vs Headless) — these require business context, cluster knowledge, and cost/reliability trade-offs. AI can list options; you decide.
3. Handling "Dirty Work"
etcd disk full, node NotReady, CoreDNS intermittent timeouts, iptables/IPVS migration pitfalls — these demand on-site experience the author has documented in prior articles.
4. Establishing Standards and Governance
Who defines YAML standards? Who writes admission webhooks for validation? Who sets resource quota policies, naming conventions, label taxonomies? This "infrastructure-level" work outweighs writing individual manifests.
How I Use AI Now
Shifted from "write YAML" to "accelerate". Four scenarios:
Scenario 1: Rapid Draft Generation
Saves ~60% time. Provide explicit parameters in the prompt:
Help me write a K8s Deployment YAML with:
- Spring Boot service, port 8080
- readinessProbe and livenessProbe using /actuator/health
- Resource limits: requests cpu 500m memory 1Gi, limits cpu 2 memory 2Gi
- Environment variables from ConfigMap for DB connection
- Run as non-root user
- PodDisruptionBudget with minAvailable 2Scenario 2: Explanation and Debugging Assistance
Paste unfamiliar CRDs or error events for instant clarification:
What does this K8s event mean?
"0/3 nodes are available: 1 node(s) had taint {node.kubernetes.io/not-ready: }, that the pod didn't tolerate."Scenario 3: Documentation and Scripts
Runbooks, oncall flows, batch shell scripts — AI handles these well.
Scenario 4: Learning New Technologies
e.g., Gateway API concepts and examples faster than reading official docs.
What Should Ops Engineers Really Worry About?
Polarization: engineers who leverage AI for higher-level work (architecture, SRE, observability, cost optimization) will thrive; those stuck in repetitive YAML tasks will be replaced. Additionally, AI lowers the barrier for developers to write their own YAML, reducing ops tickets but increasing risk of misconfigured manifests in production. The ops role shifts from "writing YAML for others" to "ensuring others' YAML works in this cluster" via tooling, platforms, standards, and automation.
My Advice
Adopt AI, don't resist. Early adoption gives an efficiency edge; resistance won't make AI disappear.
Move up the stack. Focus on architecture design, SRE practices, observability, cost optimization — areas AI cannot replace soon.
Deepen fundamental understanding. AI writes YAML but cannot explain why a Pod sticks in ContainerCreating, why Service traffic is uneven, or why rolling updates drop connections.
Build platforms and tools. Codify team experience into self-service platforms, automated YAML generation and validation. Your value becomes "enabling the whole team to operate efficiently".
Keep learning. Kubernetes evolves rapidly (Gateway API, WASM, Serverless, eBPF). AI cannot decide what to learn next.
Final thought: Don't fear AI; fear peers who master AI first. Tools don't replace people — people using tools replace those who don't.
Closing Story
The new colleague's Pod failed due to a runAsUser conflict with host directory permissions. His AI-generated YAML had runAsNonRoot: true, but he didn't understand when that setting causes issues. AI can write YAML, but it cannot help you understand your own YAML. The ops path forward: not competing on YAML speed, but on deeper understanding of what lies behind the YAML.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
