cgroups in Practice (Docker & K8s)
How container platforms translate user-facing flags to cgroup files
On this page · 8 sections
Docker Flag to cgroup File Mapping
Docker passes resource settings through its runtime to the kernel's cgroup controllers. The resulting files depend on the cgroup version, runtime, and cgroup driver; inspect the running container rather than assuming a fixed path or conversion.
CPU Mappings
--cpus 2 --> cpu.max = 200000 100000
--cpus 0.5 --> cpu.max = 50000 100000
--cpu-shares 512 --> relative CPU weight (runtime conversion on cgroup v2)
--cpuset-cpus "0,1" --> cpuset.cpus = 0,1
Memory Mappings
--memory 512m --> memory.max = 536870912 (512 * 1024 * 1024)
--memory-reservation 256m --> soft memory protection (memory.low on cgroup v2)
--memory-swap 1g --> memory.swap.max = 536870912 (1G total - 512M RAM = 512M swap)
I/O and PIDs Mappings
--pids-limit 100 --> pids.max = 100
--device-read-bps /dev/sda:10mb --> io.max = 8:0 rbps=10485760
--device-write-iops /dev/sda:1000 --> io.max = 8:0 wiops=1000
Tip
Verify it yourself: On a cgroup v2 host, run a container with --cpus 2 --memory 512m. Find its host PID with docker inspect, read /proc/<pid>/cgroup, then inspect cpu.max and memory.max under that cgroup. The path depends on the driver and configuration.
cgroup Driver: systemd vs cgroupfs
Container runtimes need to create cgroup directories and write to files. There are two approaches to managing the cgroup hierarchy:
How it works: The container runtime directly creates directories under /sys/fs/cgroup/ and writes to the files itself.
- Simple implementation
- Runtime has full control over the hierarchy
- Can conflict with systemd (both try to manage the same tree)
- systemd may "clean up" cgroups it did not create
# cgroupfs path
/sys/fs/cgroup/docker/
<container-id>/
cpu.max
memory.maxHow it works: The container runtime asks systemd to create transient scopes/slices via D-Bus API. systemd manages the hierarchy.
- Cooperates with systemd (no conflicts)
- Uses proper systemd unit management
- Kubernetes strongly recommends this driver
- kubelet and runtime must use the same driver
# systemd path
/sys/fs/cgroup/system.slice/
docker-<container-id>.scope/
cpu.max
memory.maxWarning
Match the kubelet and runtime cgroup drivers. Kubernetes recommends systemd on systemd hosts and with cgroup v2. A mismatched driver configuration can destabilize resource management under pressure.
Kubernetes Resource Model
Kubernetes exposes requests and limits. Requests guide scheduling and CPU sharing under contention; limits constrain runtime usage. Their exact cgroup mapping depends on the resource, cgroup version, runtime, and enabled Kubernetes features.
Requests vs Limits
docker run \
--cpus 2 \
--memory 512m \
my-appDocker also exposes relative CPU shares and a soft memory reservation. These are not guaranteed allocations; --cpus and --memory set ceilings.
resources:
requests:
cpu: "500m"
memory: "256Mi"
limits:
cpu: "2"
memory: "512Mi"K8s uses requests to schedule pods and weight CPU under contention. Limits set runtime ceilings; a memory request is not, by itself, a hard reserved floor.
K8s Resource to cgroup File Mapping
| K8s Resource | cgroup File | Behavior |
|---|---|---|
requests.cpu: "500m" |
cpu.weight (typically) |
Scheduling input and proportional CPU share under contention. It is not a hard CPU reservation. |
limits.cpu: "2" |
cpu.max = 200000 100000 |
Hard throttle. Process is paused when it exceeds 2 CPUs per period. |
requests.memory: "256Mi" |
Primarily a scheduling input | By default, no memory protection is set. With cgroup v2 and Memory QoS memoryReservationPolicy: TieredReservation, Kubernetes 1.37 sets memory.min for Guaranteed Pods or memory.low for Burstable Pods. The gate is beta and enabled by default, but this policy is opt-in. |
limits.memory: "512Mi" |
memory.max = 536870912 |
Hard limit. If usage cannot be reclaimed below it under memory pressure, the kernel may trigger an OOM kill. |
Warning
CPU vs Memory limits behave very differently. Exceeding a CPU limit causes throttling (the process slows down but keeps running). Exceeding a memory limit causes OOM kill (the process dies). This is why setting memory limits too low is more dangerous than setting CPU limits too low.
QoS Classes
Kubernetes classifies Pods by their CPU and memory requests and limits. For node-pressure eviction, kubelet ranks Pods by whether usage exceeds requests, Pod Priority, and usage relative to requests. QoS is a useful shorthand for likely behavior, not the direct ranking rule.
QoS classes (typical memory-pressure behavior)
BestEffort
No requests or limits set at all
Often evicted early · Lowest priority
resources: {} # nothing specifiedBurstable
requests < limits for at least one resource (or limits set without requests)
Depends on usage and Pod Priority · Middle priority
resources:
requests: { cpu: "250m", memory: "128Mi" }
limits: { cpu: "1", memory: "512Mi" }Guaranteed
requests == limits for BOTH cpu and memory (on every container in the pod)
Usually least likely to be evicted · Highest priority
resources:
requests: { cpu: "1", memory: "512Mi" }
limits: { cpu: "1", memory: "512Mi" }Note
Scheduling implication: Requests are what the scheduler uses to place pods. If a pod requests 2 CPUs and a node only has 1 CPU allocatable, the pod will not be scheduled there. Limits are enforced at runtime by cgroups. A node can be overcommitted on limits (sum of all limits > node capacity) but not on requests.
Pod cgroup Hierarchy
Kubernetes nests container cgroups beneath a Pod cgroup. Container resource settings apply per-container controls; the Pod cgroup is not automatically assigned a hard limit equal to their sum. Since Kubernetes 1.34, PodLevelResources is beta and enabled by default; when available, spec.resources can define an explicit shared Pod budget.
Guaranteed QoS pods live at this level
Pod-level cgroup (groups the pod's containers)
nginx
app
BestEffort pods — evicted first
Pod and Container Controls
- Container-level resources: Enforce each container's configured limits and provide the requests used for scheduling.
- Pod-level resources (Kubernetes 1.34+, beta, enabled by default): An explicit
spec.resourcesrequest or limit defines a shared Pod budget when the feature gate is enabled. It is not an automatic cgroup limit obtained by adding container limits. - When both levels are configured, containers retain their individual controls while sharing capacity within the pod-level budget.
ResourceQuota & LimitRange
These are admission-time controls — they are enforced by the Kubernetes API server when pods are created, not at runtime by the kernel. They determine which cgroup values kubelet will configure.
Namespace-level totals. Limits the aggregate resources across all pods in a namespace.
apiVersion: v1
kind: ResourceQuota
metadata:
name: team-quota
namespace: team-a
spec:
hard:
requests.cpu: "10"
requests.memory: "20Gi"
limits.cpu: "20"
limits.memory: "40Gi"
pods: "50"Prevents team-a from deploying more than 50 pods or requesting more than 10 CPUs total.
Per-pod/container defaults and constraints. Sets defaults and min/max per container.
apiVersion: v1
kind: LimitRange
metadata:
name: container-limits
spec:
limits:
- type: Container
default:
cpu: "500m"
memory: "256Mi"
defaultRequest:
cpu: "100m"
memory: "128Mi"
max:
cpu: "2"
memory: "1Gi"If a container does not specify requests/limits, these defaults are injected. No container can request more than 2 CPUs.
kubectl apply
User submits pod spec
API Server
LimitRange injects defaults, ResourceQuota checks totals
Scheduler
Places pod based on requests
kubelet
Creates cgroups with the computed values
Note
Key distinction: ResourceQuota rejects requests that would exceed namespace quotas; LimitRange can inject defaults or reject out-of-range values. Both act at admission, not through runtime cgroups, and neither guarantees that every resource control is configured.
HPA and VPA
Resource metrics originate from node and container accounting, then reach autoscalers through Kubernetes metrics APIs. The HPA acts on those API metrics and the requests declared for the targeted pods; it does not calculate utilization from cpu.weight.
The Metrics Pipeline
cgroup Files
cpu.stat, memory.current, etc.
kubelet
Reads cgroup files, exposes /stats/summary
metrics-server
Scrapes kubelet, aggregates into metrics API
HPA / VPA
Reads metrics, adjusts replicas or resources
Scales the number of pod replicas based on observed CPU/memory utilization.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
spec:
minReplicas: 2
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70- If average CPU > 70% of request, scale up
- If average CPU < 70% of request, scale down
- CPU utilization is current CPU usage as a percentage of the relevant container CPU requests, averaged across the targeted pods
Adjusts resource requests and limits based on observed usage history.
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
spec:
targetRef:
kind: Deployment
name: my-app
updatePolicy:
updateMode: "Recreate"
resourcePolicy:
containerPolicies:
- minAllowed:
cpu: "100m"
maxAllowed:
cpu: "4"- Analyzes usage over time (memory.current, cpu.stat)
- Recommends or sets new requests/limits
- Depending on VPA update mode and cluster support, changes can recreate the Pod or resize it in place
HPA vs VPA: When to Use Which
| Scenario | HPA | VPA |
|---|---|---|
| Stateless web servers | Ideal — add more replicas | Less useful |
| Database / stateful workloads | Hard to scale horizontally | Ideal — right-size the single instance |
| Batch jobs | Not applicable | Good for right-sizing |
| Unknown resource needs | Needs correct requests first | Can discover correct requests |
Warning
Do not use HPA and VPA on the same resource (e.g., both on CPU). They will fight: HPA tries to scale replicas based on utilization, while VPA changes the requests that define utilization. Use HPA for CPU and VPA for memory, or use one at a time.
The Full Picture: From User Config to Kernel Enforcement
Here is the complete chain from a Kubernetes manifest to kernel-level enforcement:
- User writes a pod spec with
resources.requestsandresources.limits - API server admission: LimitRange injects defaults if missing. ResourceQuota checks namespace totals. Pod is rejected or accepted.
- Scheduler places the pod on a node with enough allocatable resources (based on requests, not limits).
- kubelet on the target node asks the container runtime (containerd/CRI-O) to create the container.
- kubelet and the runtime configure cgroups (via systemd or direct cgroupfs), including applicable
cpu.max,cpu.weight,memory.max, andpids.maxvalues. On cgroup v2, memory protection from requests requires a configured Memory QoS reservation policy; it is absent by default. - Linux kernel applies the configured CPU, memory, and I/O controls during execution.
- kubelet reads cgroup accounting files (
cpu.stat,memory.current) and exposes them via/stats/summary. - metrics-server exposes resource metrics to the Kubernetes metrics API. HPA can use these metrics to change replica count; an installed VPA can recommend or apply new resource requests.
Tip
Keep the layers separate: The API declares requests and limits; scheduling and admission use them differently; kubelet and the runtime configure kernel controls. Read the live cgroup files to check what a particular node actually enforces.
References
- Kubernetes: Pod QoS and Memory QoS
- Kubernetes: requests, limits, and Pod-level resources
- Kubernetes: node-pressure eviction order
- Kubernetes: HPA resource utilization
- Linux kernel: cgroup v2 controls