Cloud Native 10 min read

Kubernetes Resource Contracts: How requests and limits Control Pod Scheduling, Throttling, and Eviction

This article explains how Kubernetes requests and limits act as a contract between workloads and the scheduler, detailing their distinct roles, the divergent consequences of CPU throttling versus memory OOMKills, the three QoS classes that determine eviction priority, and how LimitRange and ResourceQuota enforce resource policies at the namespace level.

Code Mala Tang
Code Mala Tang
Code Mala Tang
Kubernetes Resource Contracts: How requests and limits Control Pod Scheduling, Throttling, and Eviction

Two Numbers, Two Distinct Responsibilities

The article opens by reframing requests and limits not as mere configuration values but as a contract with the scheduler and the kernel. requests is the reservation the scheduler uses to place a Pod on a node that has enough unreserved capacity — not merely free capacity. A node showing 20% actual memory usage may still be unschedulable if its requests are fully booked. limits, by contrast, is the runtime ceiling enforced by Linux cgroups; it caps what the container can actually consume after it lands.

resources:
  requests:        # scheduling reservation
    cpu: 250m      # 250 millicores = 0.25 CPU
    memory: 256Mi
  limits:          # runtime ceiling
    cpu: 500m
    memory: 512Mi

Units matter: CPU uses millicores ( m, 1000m = 1 core); memory uses Mi (1024-based) versus M (1000-based) — mixing them causes subtle sizing errors.

Hitting the Ceiling: CPU Throttling vs. Memory OOMKill

The core distinction: CPU is compressible, memory is not .

CPU over limit : The kernel throttles the container — it runs slower (higher latency, stalled responses) but stays alive. Many "mystery slowdowns" trace back to CPU throttling even when utilization charts appear flat because the container is pinned at its limit.

Memory over limit : The OOM Killer terminates the container immediately — status OOMKilled — triggering a restart. A leaking workload enters a run → OOMKilled → restart → run loop.

Rule of thumb : CPU limit = slowdown (alive); memory limit = death. Give memory limits generous headroom; set CPU limits cautiously to avoid self-inflicted throttling.

QoS Tiers: Who Gets Evicted First When the Node Runs Out

Kubernetes automatically assigns every Pod a Quality of Service (QoS) class based solely on how requests and limits are set:

BestEffort — no requests or limits. First to be evicted. Suitable only for disposable tasks.

Burstable — requests set, limits higher or unset. Can burst into unused capacity; middle eviction priority. The portion above requests is reclaimed first.

Guaranteed — every container has both requests and limits and requests == limits for both CPU and memory. Highest eviction protection. Databases and core transactional services belong here.

To protect critical services , simply make every container requests == limits ; the Pod automatically becomes Guaranteed — no separate priority knob needed.

Enforcing the Contract: LimitRange and ResourceQuota

The biggest operational risk is not misconfiguration but no configuration — teams skipping resource fields entirely, flooding the cluster with BestEffort Pods that jeopardize everyone when pressure hits.

LimitRange (namespace-scoped) provides defaults for missing requests / limits and sets min/max bounds per container. Example:

apiVersion: v1
kind: LimitRange
metadata:
  name: default-limits
spec:
  limits:
  - default:
      cpu: 500m
      memory: 512Mi
    defaultRequest:
      cpu: 100m
      memory: 128Mi
    type: Container

ResourceQuota (namespace-scoped) caps the aggregate resources a namespace can request (e.g., 10 CPU cores, 20Gi memory). Creation of Pods that would exceed the quota is rejected. Essential for multi-tenant clusters to prevent one team from starving others.

Together they turn "resource as contract" from a guideline into an enforced guardrail — the platform engineering safety net.

Hard Conclusions

requests

governs scheduling (reservation, decides placement); limits governs runtime ceiling (kernel-enforced).

Scheduling looks at unreserved requests capacity , not actual free resources.

CPU over limit = throttling (slow); memory over limit = OOMKilled (dead).

QoS class is derived automatically from requests / limits settings.

Critical services gain strongest eviction resistance by setting requests == limits per container (Guaranteed).

Use LimitRange to supply defaults and bounds; use ResourceQuota to cap namespace totals — making the contract mandatory.

The next article in the series will cover the three probe types — liveness, readiness, startup — that tell the cluster whether a Pod is alive, ready for traffic, or still starting up.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Kubernetesresource managementschedulingQoSrequestsevictionLimitRangeCPU throttlinglimitsOOMKilledResourceQuota
Code Mala Tang
Written by

Code Mala Tang

Read source code together, write articles together, and enjoy spicy hot pot together.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.