Executive Summary
cgroup v2 in Kubernetes is now the default resource controller, and the change stays invisible until it is not. The kernel enforces the same limits. What moved is where the numbers live and how tools read them. Applications that consult the cgroup filesystem directly, or infer their own memory ceiling from the host, can read the wrong value and size themselves against a machine they do not own.
For teams running dense workloads, the symptom is an out of memory kill on a node with plenty of headroom in the dashboard. The fix is a version inventory rather than a config change. What follows is what breaks, and how to check for it before it pages you at 3am.
cgroup v2 reached general availability in Kubernetes 1.25 and the kubelet now detects it automatically. On a distribution that enables it by default there is no flag to set. The requirements are a kernel at 5.8 or later, a container runtime that understands the unified hierarchy, and the systemd cgroup driver on both the kubelet and the runtime. The Kubernetes documentation on cgroups is the reference to read before a node upgrade window, and containerd v1.4 plus CRI-O v1.20 are the runtime floors.
The reason it matters is that cgroup v2 replaced a set of separate controllers with one unified hierarchy and a different file layout. Software that walks the old paths finds nothing. Software that falls back to the host total finds too much. Both failure modes hand a container room it does not have, and the container spends it.
The Host Total Is the Trap, Not the Limit
The clearest case is a runtime that sizes its own heap. Node.js reads cgroup v2 memory limits through libuv starting with version 20.3.0. The 18 release line does not reliably detect them. On an affected build the process reads the node’s total memory instead of the pod limit, picks a heap to match, and walks into an out of memory kill under load. The workaround is to set the heap explicitly, for example with the max old space size flag, until the runtime is upgraded.
Java has the same shape and a longer upgrade tail. OpenJDK builds from jdk8u372, 11.0.16 and 15 onward understand cgroup v2, as do IBM Semeru builds from 8.0.382.0, 11.0.20.0 and 17.0.8.0. Older builds size the heap against the host and quietly overcommit. Go binaries that read CPU quota need uber-go/automaxprocs at v1.5.1 or higher for the same reason.
Your Monitoring and Security Agents Touch the Same Files
The workload is only half of it. Anything that reads the cgroup filesystem on the node has to speak the new layout. A standalone cAdvisor DaemonSet needs v0.43.0 or later. Third party monitoring and security agents that scrape cgroup paths need v2-capable releases, and a stale agent is how a limit that is enforced correctly still shows up wrong on a graph.
This is the part operators miss, because the kubelet migration looks complete. The kubelet is fine. The node exporter is fine. The one agent nobody versioned is the one misreporting, and it reports a healthy node while the scheduler’s own accounting is unchanged.

Three Checks Before Your Next Node Upgrade
Ask which of your runtimes size themselves from the environment, and pin the heap for any that predate cgroup v2 support. Ask whether every node-level agent has a v2-capable release in your image. And ask what your pod density would look like if the memory limits were finally read correctly, which is often tighter than the current schedule assumes.
The migration is a version problem wearing a config problem’s clothes. Read the limits where the kernel put them, update the agents that read them, and the cutover stays quiet. Related reading on how memory pressure changes packing is in our piece on Kubernetes node swap and agent density.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.
