Controllers (CPU, Memory, I/O)
CPU throttling, memory limits, I/O control, and fork bomb protection
On this page · 6 sections
CPU Controller
The CPU controller manages how much processor time a cgroup gets. It has two distinct mechanisms: proportional sharing (how to divide CPU when contended) and bandwidth limiting (hard cap regardless of contention).
cpu.weight — Proportional CPU Sharing
cpu.weight (v2) / cpu.shares (v1)
- Default: 100 (range: 1 - 10000)
- Only matters when CPU is contended — if only one cgroup is busy, it gets 100% regardless of its weight
- Weights are relative: among equally runnable fair-class cgroups, weight 200 gives roughly 2x the CPU share of weight 100 under sustained contention
- Think of it as "priority shares" — the kernel divides available CPU proportionally
With one CPU and three equally runnable fair-class cgroups, weights give these approximate shares under sustained contention:
| Cgroup | Weight | Share |
|---|---|---|
| App A | 200 | 50% |
| App B | 100 | 25% |
| App C | 100 | 25% |
If B and C become idle, A can use the available CPU; weights govern competition, not a hard quota.
cpu.max — Hard CPU Bandwidth Limit
cpu.max (v2) / cpu.cfs_quota_us + cpu.cfs_period_us (v1)
- Format:
$MAX $PERIOD(both in microseconds) - The cgroup can use at most
$MAXmicroseconds of CPU per$PERIODmicrosecond window - Default:
max 100000(unlimited quota, 100ms period) - CFS bandwidth throttling: when quota is exhausted, all threads in the cgroup are paused until the next period
| cpu.max Value | Effective Limit | Docker Equivalent |
|---|---|---|
200000 100000 |
2 CPUs (200ms per 100ms period) | --cpus 2 |
100000 100000 |
1 CPU (100ms per 100ms period) | --cpus 1 |
50000 100000 |
0.5 CPUs (50ms per 100ms period) | --cpus 0.5 |
25000 100000 |
0.25 CPUs (25ms per 100ms period) | --cpus 0.25 |
max 100000 |
Unlimited | (default — no --cpus flag) |
Note
The period is typically 100ms (100000us). With 200000 100000, the cgroup can burn 200ms of CPU time per 100ms wall-clock period. Since this exceeds 100ms, it must use multiple CPU cores to fill the quota — that is how you express "2 CPUs" in cgroup terms.
cpuset Controller — CPU Pinning
cpuset: Pin Processes to Specific Cores
cpuset.cpus=0-3means only use cores 0, 1, 2, 3cpuset.cpus=0,2,4means only use cores 0, 2, and 4cpuset.mems= NUMA node pinning (which memory banks to use)
Use Cases for CPU Pinning
- Latency-sensitive workloads (avoid cache thrashing from core migration)
- NUMA-aware placement (keep processes close to their memory)
- Isolating noisy neighbors (dedicate cores to critical services)
- Real-time workloads that need consistent performance
cpuset vs cpu.max
- cpuset: controls which cores — spatial partitioning
- cpu.max: controls how much time — temporal throttling
- They can be combined: pin to cores 0-3, AND limit to 2 CPUs of bandwidth
- cpuset alone does not limit CPU usage within the allowed cores
Memory Controller
The memory controller manages RAM and swap consumption. It provides a layered set of limits — from hard kills to soft pressure — giving fine-grained control over how the kernel handles memory scarcity.
Memory Limit Hierarchy
| Control / example | Effect |
|---|---|
memory.max512 MiB | Hard ceiling; cgroup OOM if reclaim cannot bring usage below it. |
memory.high400 MiB | Throttling and heavy reclaim above this boundary, without directly invoking OOM. |
memory.low128 MiB | Best-effort reclaim protection below the effective boundary. |
memory.min64 MiB | Hard reclaim protection below the effective boundary, limited by ancestors. |
These values illustrate separate controls, not an automatically configured tier. Limits set ceilings; protection affects reclaim.
OOM Killer Behavior
What Happens When memory.max Is Hit
- The kernel attempts memory reclaim (swap, page cache eviction) first
- If reclaim cannot free enough memory, the OOM killer activates
- The OOM killer selects and kills a process within the cgroup — not system-wide
memory.oom.group = 1— treats the cgroup and descendants as one workload for cgroup OOM kills, except tasks withoom_score_adj=-1000- A Kubernetes
OOMKilledcontainer can result from such a memory limit; inspect the container and node events to identify the cause.
The OOM killer may pick one process in the cgroup. Other processes can continue running.
May leave the application in an inconsistent state if only one worker dies.
The cgroup and descendants are treated as one workload for a cgroup OOM kill, apart from tasks with oom_score_adj=-1000.
This avoids leaving most of a workload running after one task is killed.
Swap Limits and Accounting
| File | Purpose |
|---|---|
memory.swap.max |
Maximum swap usage (independent of RAM limit). Set to 0 to disable swap for this cgroup. |
memory.current |
Current total memory usage (RSS + page cache) |
memory.swap.current |
Current swap usage |
memory.stat |
Detailed breakdown: anon (RSS), file (page cache), kernel stacks, slab, sock, shmem, swap |
memory.peak |
High-water mark — maximum memory.current ever reached |
$ cat /sys/fs/cgroup/my-app/worker/memory.stat
anon 25165824
file 16777216
kernel 1048576
sock 0
shmem 0
zswap 0
zswapped 0
file_mapped 4194304
file_dirty 0
file_writeback 0
swapcached 0
anon_thp 0
inactive_anon 12582912
active_anon 12582912
inactive_file 8388608
active_file 8388608
pgfault 15234
pgmajfault 3
...I/O Controller
The I/O controller limits disk read/write bandwidth and IOPS per block device. Like the CPU controller, it supports both proportional sharing and hard limits.
I/O Control Files
| File | Purpose | Format |
|---|---|---|
io.weight |
Proportional I/O bandwidth (like cpu.weight) | default 100 or MAJ:MIN 200 |
io.max |
Hard I/O limits per device | MAJ:MIN rbps=X wbps=Y riops=Z wiops=W |
io.latency |
Latency-based I/O control | MAJ:MIN target=Xus |
io.stat |
Per-device I/O statistics | (read-only) |
io.max Examples
# # Find the major:minor numbers for /dev/sda
$ lsblk -o NAME,MAJ:MIN
NAME MAJ:MIN
sda 8:0
|-sda1 8:1
|-sda2 8:2
# # Limit /dev/sda reads to 10 MB/s
# echo "8:0 rbps=10485760" > /sys/fs/cgroup/my-app/worker/io.max
# # Limit both reads and writes, plus IOPS
# echo "8:0 rbps=10485760 wbps=5242880 riops=1000 wiops=500" > io.max
# # Limit format explanation:
# # rbps = read bytes per second (10 MB/s)
# # wbps = write bytes per second (5 MB/s)
# # riops = read I/O ops per second (1000)
# # wiops = write I/O ops per second (500)Warning
Caveat: I/O cgroup throttling only works reliably with direct I/O or when the filesystem supports cgroup-aware writeback. With buffered I/O, writes go to the page cache first and are written back by kernel threads that may be in a different cgroup — the limits may not apply as expected. ext4 and btrfs support cgroup writeback; XFS gained support in kernel 5.18+.
PIDs Controller
The simplest controller — it limits the maximum number of processes (and threads) in a cgroup.
Fork Bomb Protection
pids.max— maximum number of processes/threads allowed in the cgrouppids.current— current count of processes/threads- When the limit is hit,
fork()andclone()returnEAGAIN - Without this, a runaway container running
:(){ :|:& };:can exhaust the host's entire PID space and freeze the system
# Fork bomb in a container
:(){ :|:& };:
# Result: exponential process creation
# Consumes all PIDs on the host
# Host becomes unresponsive
# Even SSH may stop working# Fork bomb in a container
:(){ :|:& };:
# Result: stops at 100 processes
# fork() returns EAGAIN
# Container is broken, but host is fine
# Other containers unaffectedDevices Controller
Device access can be restricted per cgroup. Container runtimes such as Docker configure a policy so ordinary containers cannot open arbitrary host devices.
Device Access Model
- v1: Explicit device controller with
devices.allowanddevices.deny - v2: Device access is managed via eBPF programs attached to the cgroup (the
devicescontroller was replaced byBPF_PROG_TYPE_CGROUP_DEVICE) - Policy comes from the runtime: Docker restricts device access by default and allows selected devices; cgroup v2 itself does not impose a deny-all BPF policy
Typical Container Device Allowlist
| Device | Type | Allowed | Purpose |
|---|---|---|---|
/dev/null |
char 1:3 | Yes | Discard output |
/dev/zero |
char 1:5 | Yes | Zero bytes source |
/dev/urandom |
char 1:9 | Yes | Random number generator |
/dev/tty |
char 5:0 | Yes | Terminal access |
/dev/sda |
block 8:0 | No | Host disk — denied |
/dev/kmsg |
char 1:11 | No | Kernel message buffer — denied |
Tip
Docker --device flag: When you run docker run --device /dev/fuse, Docker adds that device to the cgroup's allowlist and bind-mounts it into the container's filesystem. Without --device, the container cannot access it even if it has the right capabilities.
Controller Summary
| Controller | Key Files (v2) | What It Controls | Default Behavior |
|---|---|---|---|
| cpu | cpu.weight, cpu.max |
CPU time (proportional sharing + bandwidth) | Unlimited, fair share |
| memory | memory.max, memory.high, memory.low |
Memory usage; swap has separate memory.swap.* controls |
No memory ceiling by default |
| io | io.weight, io.max |
Disk bandwidth and IOPS | Unlimited |
| pids | pids.max |
Process/thread count | Unlimited |
| cpuset | cpuset.cpus, cpuset.mems |
CPU core and NUMA pinning | All CPUs, all NUMA nodes |
| devices | (eBPF-based in v2) | Block/char device access | Access depends on attached device BPF programs |
Note
Defaults differ by control: Hard limits such as memory.max and pids.max default to max; relative weights have numeric defaults. Ancestor cgroups and the runtime's configuration also constrain what a container can use.