| in blog | Kubernetes Blog |
|---|---|
| published date | 2026-10-06 |
| original entry | The Shift to cgroup v2 in Kubernetes: What You Need to Know |
In Linux, cgroups (control groups) are a kernel feature used for managing system resources. Kubernetes uses cgroups to allocate resources like CPU and memory to containers, ensuring that applications run smoothly without interfering with each other. With the release of Kubernetes v1.31, support for v1 cgroup management moved into maintenance mode. Support for v2 cgroup management has been stable since Kubernetes v1.25.
Compared with cgroup v1, cgroup v2 provides a single unified hierarchy, a more consistent interface, and a stronger foundation for resource isolation and modern resource-management features.
Kubernetes has deprecated cgroup v1.
Starting with Kubernetes v1.35, failCgroupV1 defaults to true, so the kubelet does not start
on a cgroup v1 node by default. Administrators can temporarily set failCgroupV1: false in the
kubelet configuration file, but removal will
follow the Kubernetes deprecation policy.
Further removal work is tracked in KEP-5573: Remove cgroup v1 support.
If you are still on a release older than v1.35, migrate every Linux node to
cgroup v2 before upgrading, or plan to set the temporary failCgroupV1: false
override. If you are already on v1.35 or later, confirm that every Linux node
runs cgroup v2 (or that you intentionally keep the override). Under the default
configuration, a remaining cgroup v1 node fails during kubelet startup.
For kubeadm-managed clusters, Kubernetes v1.35 also makes this an earlier, stricter check. The
SystemVerification preflight check, provided by k8s.io/system-validators, returns an error during
kubeadm init, kubeadm join, and kubeadm upgrade when it detects cgroup v1 with kubelet v1.35 or
later; with an older kubelet, the check remains a warning. See
kubernetes/system-validators#1.12.1 release notes for details.
The top FAQs cover three main areas: why to migrate, the benefits and drawbacks, and key points to keep in mind when using cgroup v2.
The Linux kernel documentation describes both interfaces:
Let's enumerate some known issues.
active_file memory is not considered available memoryThe kubelet treats active_file memory as not reclaimable. For I/O-intensive workloads, a large
page cache can therefore make the kubelet report memory pressure and evict Pods. This is a
known kubelet issue
(kubernetes/kubernetes#43916); migrating to
cgroup v2 does not by itself change that calculation. The documented workaround is to set equal
memory requests and limits for containers that perform intensive I/O, after measuring an appropriate
value.
Memory QoS was introduced as an alpha feature in Kubernetes v1.22 and updated in v1.27. It remains alpha in v1.36, but now separates memory throttling from memory reservation and adds tiered memory protection:
Memory QoS is available only on Linux nodes that use cgroup v2. It relies on the cgroup v2 memory
controller: memory.high provides throttling, while memory.min and memory.low provide hard and
soft protection when tiered reservation is enabled. cgroup v1 cannot provide this protection model.
Enabling the MemoryQoS feature gate applies memory.high throttling to Burstable containers.
The threshold is derived from the request, limit, and memoryThrottlingFactor (default 0.9).
memoryReservationPolicy: None is the default. It does not write memory.min or memory.low.
memoryReservationPolicy: TieredReservation maps Guaranteed Pod memory requests to memory.min
(hard protection) and Burstable Pod requests to memory.low (soft protection). BestEffort Pods
receive neither protection.
The kubelet exposes Alpha metrics for the total memory.min and memory.low reservations on a node.
Kernel 5.9 or later is recommended. On older kernels, memory.high reclaim can trigger a known
livelock; from v1.36 the kubelet logs a warning when Memory QoS is enabled on an affected kernel.
For example, to opt in to tiered protection:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
featureGates:
MemoryQoS: true
memoryReservationPolicy: TieredReservation
memoryThrottlingFactor: 0.9
The overall Kubernetes recommendation is not to enable Alpha features in production; however, if you
judge that memory QoS with tiered reservations is useful for your platform in production, make sure
to test the configuration and account for hard-reserved memory before you enable
the TieredReservation feature gate.
See the
Pod QoS documentation
for the current mapping of Kubernetes QoS classes to cgroup v2 controls.
On cgroup v2 nodes, the kubelet defaults singleProcessOOMKill to false. It therefore sets
memory.oom.group for each container cgroup so that an OOM event kills all processes in that
container together, rather than leaving a partially functioning multi-process container. Set
singleProcessOOMKill: true only if you need the cgroup v1-compatible behavior where the kernel
may kill one process at a time. See the
KubeletConfiguration reference
for this setting.
This behavior is scoped to a container cgroup, not the whole Pod. Also, cgroup.kill is a separate
administrative interface: writing 1 to it sends SIGKILL to every process in that cgroup and its
descendants; it does not configure OOM behavior. The cgroup v2 memory controller additionally
provides memory.events counters that monitoring systems and userspace OOM managers can observe.
In cgroup v1, delegating controllers to less privileged containers may be dangerous.
Unlike cgroup v1, cgroup v2 officially supports delegation. Most implementations of rootless containers rely on systemd for delegating v2 controllers to non-root users.
This delegation mechanism is separate from Kubernetes Pod user namespaces, which map container users to unprivileged host users. Pod user namespace support graduated to stable in Kubernetes v1.36; check its filesystem, kernel, CRI runtime, and OCI runtime prerequisites before enabling it.
/run/cilium/cgroupv2.KubeletPSI is stable and locked on). PSI
requires cgroup v2, Linux 4.20 or later, CONFIG_PSI=y, and a kernel not booted with
psi=0. The kubelet surfaces the data through the
Summary API and
/metrics/cadvisor.automaxprocs versions.cgroup v1 uses cpu.shares, whereas cgroup v2 uses cpu.weight. Newer OCI runtimes use an improved
non-linear conversion that preserves the default priority and gives small CPU requests more usable
granularity. The change is implemented in the OCI runtime rather than Kubernetes: it is available
in crun v1.23 and runc v1.3.2. After upgrading a runtime, monitoring or policy tools that predict
exact cpu.weight values may need updates. Read
New Conversion from cgroup v1 CPU Shares to v2 CPU Weight
for the formula, examples, and compatibility considerations.
In-place Pod vertical scaling graduated to stable in Kubernetes v1.35. Kubernetes v1.36 then [enabled] (/blog/2026/04/30/kubernetes-v1-36-inplace-pod-level-resources-beta/) in-place vertical scaling for Pod-level resources by default, as a Beta feature. The kubelet coordinates changes between the Pod-level and container cgroups so that increases create headroom before container limits grow, while decreases constrain containers before shrinking the Pod-level boundary. Accurate aggregate enforcement for this v1.36 feature requires cgroup v2.
Here's what you need to use cgroup v2 with Kubernetes. First up, you need to be using a version of Kubernetes with support for v2 cgroup management; that's been stable since Kubernetes v1.25 and all supported Kubernetes releases include this support.
For now, you can opt back in to use cgroup v1; the Kubernetes project recommends using cgroup v2, but in Kubernetes 1.36 (the current release) the cgroup v1 option remains supported as a fallback. That fallback is scheduled for removal in Kubernetes v1.38. If you are running an older cluster, plan to migrate; if you are setting up a new cluster with Linux nodes, you should prefer cgroup v2. In either case, review both the Kubernetes runtime documentation and (if relevant) the containerd compatibility matrix.
When Kubernetes was first announced, in 2014, only v1 cgroup existed. Version 2 cgroup management first appeared in Linux kernel 4.5, released in 2016.
io, memory, and pids controllers were supported.cpu controller.cpu.stat file was added in Linux 5.8.memory.high livelock fix used by Memory QoS is present in Linux 5.9 and later.memory.peak was added in Linux 5.19.Configure the kubelet's cgroup driver to match the container runtime cgroup driver.
If you use kubeadm to manage your cluster, Kubernetes recommends that you use the systemd cgroup driver,
because kubeadm manages the kubelet as a systemd service. For other management tooling,
check the documentation for the tool you're using to manage your cluster.
If you can pick either option, I recommend using the systemd driver.
Whatever tooling you've chosen, the kubelet automatically tries to detect the runtime's recommended cgroup driver.
This automatic detection relies on using a runtime
that implements the RuntimeConfig
CRI RPC (for example: containerd v2.0+ or CRI-O v1.28+).
If you're using a container runtime that supports cgroup v2 but doesn't support automatic cgroup driver detection, you can manually configure an override by editing the kubelet configuration file. For example:
apiVersion: kubelet.config.k8s.io/v1beta1
kind: KubeletConfiguration
cgroupDriver: systemd
Automatic discovery of the runtime's cgroup driver through the CRI, tracked by
KEP-4033, graduated to stable in Kubernetes v1.34. It requires a runtime
that implements the RuntimeConfig CRI RPC (containerd v2.0+ or CRI-O v1.28+). When available,
the kubelet uses the value reported by the runtime instead of its configured cgroupDriver value.
Tools and commands that you should know about cgroups:
stat -fc %T /sys/fs/cgroup/: Check whether cgroup v2 is enabled; it returns cgroup2fs.systemctl list-units 'kube*' --type=slice or --type=scope: List Kubernetes-related units
that systemd currently has in memory.bpftool cgroup list /sys/fs/cgroup/*: List all programs attached to the cgroup CGROUP.systemd-cgls /sys/fs/cgroup/*: Recursively show control group contents.systemd-cgtop: Show top control groups by their resource usage.tree -L 2 -d /sys/fs/cgroup/kubepods.slice: Show Pods' related cgroups directories.Work from the API object down to the node. Identify the node, then compare desired resources in the spec with the enacted values in status (in-place resize can leave those out of sync):
kubectl get pod <pod-name> -n <namespace> \
-o jsonpath='{.spec.nodeName}{"\n"}'
kubectl get pod <pod-name> -n <namespace> \
-o jsonpath='{range .spec.containers[*]}{.name}{" spec\t"}{.resources}{"\n"}{end}'
kubectl get pod <pod-name> -n <namespace> \
-o jsonpath='{range .status.containerStatuses[*]}{.name}{" status\t"}{.resources}{"\n"}{end}'
On that node, as root, with the systemd cgroup driver and cgroup v2 (these
snippets also need jq). Pin the container with kubelet labels; a container
name alone is not unique on a node. Use crictl ps -a if the container is
not running:
CONTAINER_ID=$(crictl ps \
--label io.kubernetes.pod.namespace=<namespace> \
--label io.kubernetes.pod.name=<pod-name> \
--name <container-name> -q | head -n1)
# OCI view (not the CRI protobuf field names)
crictl inspect "$CONTAINER_ID" | jq '.info.runtimeSpec.linux.resources'
# Kernel cgroup path for this container
PID=$(crictl inspect "$CONTAINER_ID" | jq -r '.info.pid')
CGROUP="/sys/fs/cgroup$(awk -F: '$1=="0"{print $3}' /proc/$PID/cgroup)"
cat "$CGROUP/cpu.weight" # request.cpu → shares → weight; not a limit
cat "$CGROUP/cpu.max" # limit.cpu as "quota period"; unlimited is "max <period>"
cat "$CGROUP/memory.max" # limit.memory; unlimited is "max"
# Pod-level cgroup (parent slice), used by Pod-level resources
cat "$(dirname "$CGROUP")/cpu.max" "$(dirname "$CGROUP")/memory.max"
# Present when MemoryQoS is enabled
cat "$CGROUP/memory.high" "$CGROUP/memory.min" "$CGROUP/memory.low"
SCOPE=$(basename "$CGROUP")
systemctl show "$SCOPE" \
-p CPUWeight -p CPUQuotaPerSecUSec -p CPUQuotaPeriodUSec -p MemoryMax
Do not treat info.runtimeSpec.linux.cgroupsPath as a filesystem path when
the runtime uses the systemd driver; that value is a systemd unit path
(slice:runtime:id), not a directory under /sys/fs/cgroup.
Expected mapping at each layer:
spec.containers[*].resources (desired) and
status.containerStatuses[*].resources (enacted after in-place resize).
Pod-level resources are on spec.resources and the parent pod slice.cpu_shares, cpu_quota, cpu_period, memory_limit_in_bytes, plus
unified for Memory QoS. crictl inspect does not show these names; it
shows the OCI spec the runtime derived from them.linux.resources.cpu.{shares,quota,period},
linux.resources.memory.limit, and linux.resources.unified. The runtime
converts shares, quota, and period into cgroup v2 files. After upgrading
crun or runc, the shares-to-weight formula may change; see
CPU weight conversion.CPUWeight, CPUQuotaPerSecUSec, CPUQuotaPeriodUSec,
MemoryMax. Unlimited or unset values often appear as [not set] or
infinity.cpu.weight (from CPU request; an unset request still yields the
default of 2 shares), cpu.max, and memory.max on the container scope.
Memory QoS adds memory.high, memory.min, and memory.low.