Skip to content
Menu

Virtualization8 min read

cgroups in Practice (Docker & K8s)

How container platforms translate user-facing flags to cgroup files

On this page · 8 sections

Docker Flag to cgroup File Mapping

Docker passes resource settings through its runtime to the kernel's cgroup controllers. The resulting files depend on the cgroup version, runtime, and cgroup driver; inspect the running container rather than assuming a fixed path or conversion.

CPU Mappings

--cpus 2 --> cpu.max = 200000 100000

--cpus 0.5 --> cpu.max = 50000 100000

--cpu-shares 512 --> relative CPU weight (runtime conversion on cgroup v2)

--cpuset-cpus "0,1" --> cpuset.cpus = 0,1

Memory Mappings

--memory 512m --> memory.max = 536870912 (512 * 1024 * 1024)

--memory-reservation 256m --> soft memory protection (memory.low on cgroup v2)

--memory-swap 1g --> memory.swap.max = 536870912 (1G total - 512M RAM = 512M swap)

I/O and PIDs Mappings

--pids-limit 100 --> pids.max = 100

--device-read-bps /dev/sda:10mb --> io.max = 8:0 rbps=10485760

--device-write-iops /dev/sda:1000 --> io.max = 8:0 wiops=1000

Tip

Verify it yourself: On a cgroup v2 host, run a container with --cpus 2 --memory 512m. Find its host PID with docker inspect, read /proc/<pid>/cgroup, then inspect cpu.max and memory.max under that cgroup. The path depends on the driver and configuration.

cgroup Driver: systemd vs cgroupfs

Container runtimes need to create cgroup directories and write to files. There are two approaches to managing the cgroup hierarchy:

cgroupfs Driver

How it works: The container runtime directly creates directories under /sys/fs/cgroup/ and writes to the files itself.

  • Simple implementation
  • Runtime has full control over the hierarchy
  • Can conflict with systemd (both try to manage the same tree)
  • systemd may "clean up" cgroups it did not create
console
# cgroupfs path
/sys/fs/cgroup/docker/
    <container-id>/
        cpu.max
        memory.max
systemd Driver

How it works: The container runtime asks systemd to create transient scopes/slices via D-Bus API. systemd manages the hierarchy.

  • Cooperates with systemd (no conflicts)
  • Uses proper systemd unit management
  • Kubernetes strongly recommends this driver
  • kubelet and runtime must use the same driver
console
# systemd path
/sys/fs/cgroup/system.slice/
    docker-<container-id>.scope/
        cpu.max
        memory.max

Warning

Match the kubelet and runtime cgroup drivers. Kubernetes recommends systemd on systemd hosts and with cgroup v2. A mismatched driver configuration can destabilize resource management under pressure.

Kubernetes Resource Model

Kubernetes exposes requests and limits. Requests guide scheduling and CPU sharing under contention; limits constrain runtime usage. Their exact cgroup mapping depends on the resource, cgroup version, runtime, and enabled Kubernetes features.

Requests vs Limits

Docker
bash
docker run \
  --cpus 2 \
  --memory 512m \
  my-app

Docker also exposes relative CPU shares and a soft memory reservation. These are not guaranteed allocations; --cpus and --memory set ceilings.

Kubernetes
yaml
resources:
  requests:
    cpu: "500m"
    memory: "256Mi"
  limits:
    cpu: "2"
    memory: "512Mi"

K8s uses requests to schedule pods and weight CPU under contention. Limits set runtime ceilings; a memory request is not, by itself, a hard reserved floor.

K8s Resource to cgroup File Mapping

K8s Resource cgroup File Behavior
requests.cpu: "500m" cpu.weight (typically) Scheduling input and proportional CPU share under contention. It is not a hard CPU reservation.
limits.cpu: "2" cpu.max = 200000 100000 Hard throttle. Process is paused when it exceeds 2 CPUs per period.
requests.memory: "256Mi" Primarily a scheduling input By default, no memory protection is set. With cgroup v2 and Memory QoS memoryReservationPolicy: TieredReservation, Kubernetes 1.37 sets memory.min for Guaranteed Pods or memory.low for Burstable Pods. The gate is beta and enabled by default, but this policy is opt-in.
limits.memory: "512Mi" memory.max = 536870912 Hard limit. If usage cannot be reclaimed below it under memory pressure, the kernel may trigger an OOM kill.

Warning

CPU vs Memory limits behave very differently. Exceeding a CPU limit causes throttling (the process slows down but keeps running). Exceeding a memory limit causes OOM kill (the process dies). This is why setting memory limits too low is more dangerous than setting CPU limits too low.

QoS Classes

Kubernetes classifies Pods by their CPU and memory requests and limits. For node-pressure eviction, kubelet ranks Pods by whether usage exceeds requests, Pod Priority, and usage relative to requests. QoS is a useful shorthand for likely behavior, not the direct ranking rule.

QoS classes (typical memory-pressure behavior)

BestEffort

No requests or limits set at all

Often evicted early · Lowest priority

yaml
resources: {}  # nothing specified

Burstable

requests < limits for at least one resource (or limits set without requests)

Depends on usage and Pod Priority · Middle priority

yaml
resources:
  requests: { cpu: "250m", memory: "128Mi" }
  limits:   { cpu: "1",    memory: "512Mi" }

Guaranteed

requests == limits for BOTH cpu and memory (on every container in the pod)

Usually least likely to be evicted · Highest priority

yaml
resources:
  requests: { cpu: "1", memory: "512Mi" }
  limits:   { cpu: "1", memory: "512Mi" }

Note

Scheduling implication: Requests are what the scheduler uses to place pods. If a pod requests 2 CPUs and a node only has 1 CPU allocatable, the pod will not be scheduled there. Limits are enforced at runtime by cgroups. A node can be overcommitted on limits (sum of all limits > node capacity) but not on requests.

Pod cgroup Hierarchy

Kubernetes nests container cgroups beneath a Pod cgroup. Container resource settings apply per-container controls; the Pod cgroup is not automatically assigned a hard limit equal to their sum. Since Kubernetes 1.34, PodLevelResources is beta and enabled by default; when available, spec.resources can define an explicit shared Pod budget.

/sys/fs/cgroup/kubepods.slice/
kubepods-guaranteed (implied)

Guaranteed QoS pods live at this level

kubepods-burstable.slice/
kubepods-burstable-pod<uid>.slice/

Pod-level cgroup (groups the pod's containers)

<container-1-id>.scope

nginx

<container-2-id>.scope

app

kubepods-besteffort.slice/

BestEffort pods — evicted first

Pod and Container Controls

  • Container-level resources: Enforce each container's configured limits and provide the requests used for scheduling.
  • Pod-level resources (Kubernetes 1.34+, beta, enabled by default): An explicit spec.resources request or limit defines a shared Pod budget when the feature gate is enabled. It is not an automatic cgroup limit obtained by adding container limits.
  • When both levels are configured, containers retain their individual controls while sharing capacity within the pod-level budget.

ResourceQuota & LimitRange

These are admission-time controls — they are enforced by the Kubernetes API server when pods are created, not at runtime by the kernel. They determine which cgroup values kubelet will configure.

ResourceQuota

Namespace-level totals. Limits the aggregate resources across all pods in a namespace.

yaml
apiVersion: v1
kind: ResourceQuota
metadata:
  name: team-quota
  namespace: team-a
spec:
  hard:
    requests.cpu: "10"
    requests.memory: "20Gi"
    limits.cpu: "20"
    limits.memory: "40Gi"
    pods: "50"

Prevents team-a from deploying more than 50 pods or requesting more than 10 CPUs total.

LimitRange

Per-pod/container defaults and constraints. Sets defaults and min/max per container.

yaml
apiVersion: v1
kind: LimitRange
metadata:
  name: container-limits
spec:
  limits:
  - type: Container
    default:
      cpu: "500m"
      memory: "256Mi"
    defaultRequest:
      cpu: "100m"
      memory: "128Mi"
    max:
      cpu: "2"
      memory: "1Gi"

If a container does not specify requests/limits, these defaults are injected. No container can request more than 2 CPUs.

  1. kubectl apply

    User submits pod spec

  2. API Server

    LimitRange injects defaults, ResourceQuota checks totals

  3. Scheduler

    Places pod based on requests

  4. kubelet

    Creates cgroups with the computed values

Note

Key distinction: ResourceQuota rejects requests that would exceed namespace quotas; LimitRange can inject defaults or reject out-of-range values. Both act at admission, not through runtime cgroups, and neither guarantees that every resource control is configured.

HPA and VPA

Resource metrics originate from node and container accounting, then reach autoscalers through Kubernetes metrics APIs. The HPA acts on those API metrics and the requests declared for the targeted pods; it does not calculate utilization from cpu.weight.

The Metrics Pipeline

  1. cgroup Files

    cpu.stat, memory.current, etc.

  2. kubelet

    Reads cgroup files, exposes /stats/summary

  3. metrics-server

    Scrapes kubelet, aggregates into metrics API

  4. HPA / VPA

    Reads metrics, adjusts replicas or resources

HPA (Horizontal Pod Autoscaler)

Scales the number of pod replicas based on observed CPU/memory utilization.

yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
  minReplicas: 2
  maxReplicas: 10
  metrics:
  - type: Resource
    resource:
      name: cpu
      target:
        type: Utilization
        averageUtilization: 70
  • If average CPU > 70% of request, scale up
  • If average CPU < 70% of request, scale down
  • CPU utilization is current CPU usage as a percentage of the relevant container CPU requests, averaged across the targeted pods
VPA (Vertical Pod Autoscaler)

Adjusts resource requests and limits based on observed usage history.

yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
spec:
  targetRef:
    kind: Deployment
    name: my-app
  updatePolicy:
    updateMode: "Recreate"
  resourcePolicy:
    containerPolicies:
    - minAllowed:
        cpu: "100m"
      maxAllowed:
        cpu: "4"
  • Analyzes usage over time (memory.current, cpu.stat)
  • Recommends or sets new requests/limits
  • Depending on VPA update mode and cluster support, changes can recreate the Pod or resize it in place

HPA vs VPA: When to Use Which

Scenario HPA VPA
Stateless web servers Ideal — add more replicas Less useful
Database / stateful workloads Hard to scale horizontally Ideal — right-size the single instance
Batch jobs Not applicable Good for right-sizing
Unknown resource needs Needs correct requests first Can discover correct requests

Warning

Do not use HPA and VPA on the same resource (e.g., both on CPU). They will fight: HPA tries to scale replicas based on utilization, while VPA changes the requests that define utilization. Use HPA for CPU and VPA for memory, or use one at a time.

The Full Picture: From User Config to Kernel Enforcement

Here is the complete chain from a Kubernetes manifest to kernel-level enforcement:

  1. User writes a pod spec with resources.requests and resources.limits
  2. API server admission: LimitRange injects defaults if missing. ResourceQuota checks namespace totals. Pod is rejected or accepted.
  3. Scheduler places the pod on a node with enough allocatable resources (based on requests, not limits).
  4. kubelet on the target node asks the container runtime (containerd/CRI-O) to create the container.
  5. kubelet and the runtime configure cgroups (via systemd or direct cgroupfs), including applicable cpu.max, cpu.weight, memory.max, and pids.max values. On cgroup v2, memory protection from requests requires a configured Memory QoS reservation policy; it is absent by default.
  6. Linux kernel applies the configured CPU, memory, and I/O controls during execution.
  7. kubelet reads cgroup accounting files (cpu.stat, memory.current) and exposes them via /stats/summary.
  8. metrics-server exposes resource metrics to the Kubernetes metrics API. HPA can use these metrics to change replica count; an installed VPA can recommend or apply new resource requests.

Tip

Keep the layers separate: The API declares requests and limits; scheduling and admission use them differently; kubelet and the runtime configure kernel controls. Read the live cgroup files to check what a particular node actually enforces.

References

Solidnines — solidnines.com