A person working at a desk in a dark room. Photo by kartik programmer on Unsplash. Source: https://unsplash.com/photos/a-person-sitting-at-a-desk-in-a-dark-room-JrEHwkPIurA (Unsplash License).

Executive Summary

DepthFirst published a working container escape on September 22. It runs with nothing more than default Docker or Kubernetes settings. The flaw is a use-after-free in the Linux kernel code that passes file descriptors between local processes, and an unprivileged process inside a container can walk it to root on the host.

The patch has existed upstream since August 6. Ubuntu still lists its two current LTS releases as vulnerable and the fix as work in progress. That gap is the story. Containers were always a boundary of convenience rather than a wall, and AI-assisted bug hunting has collapsed the cost of finding the kernel flaws that turn one into a host compromise. The answer that holds is a workload with its own kernel.

Diagram comparing container isolation and microVM isolation. Containers share one host kernel, a microVM gives each workload its own.
One shared kernel is the containment problem. A per-workload kernel is the fix.

A container is not a small virtual machine. It is a process with a better view of the walls. It shares one kernel with the host and with every other container on the node, so a bug in that kernel is a bug in the boundary.

CVE-2026-80521 is that bug. DepthFirst found it with an in-house model built for vulnerability detection, proved it inside Google’s kernelCTF on July 24, and published the research and the exploit on September 22.

The bug lives in the code that cleans up socket references

Linux keeps local chatter between processes on the same host inside the AF_UNIX subsystem. When two processes pass an open file between them, the kernel records the reference so it can clean up circular holds later. DepthFirst found a race in that cleanup.

A tracked socket can be freed without being unlinked from a cached ring that a later cleanup pass walks. That pass then reads memory the kernel already released. It is a use-after-free, and this one is reachable from inside an ordinary container.

The same path touches most kernel-level sandboxes, including nsjail, Firejail and Bubblewrap, not just container runtimes. DepthFirst reported it on August 5. Kyle Zeng of OpenAI reported the same flaw independently. The upstream fix landed the next day.

Default settings are the exposure, not a misconfiguration

The exploit needs AF_UNIX sockets, descriptor passing, and one capability that Docker grants by default. It does not need a privileged container, a kernel module, or an unusual flag. Kubernetes baseline pod security already permits the pieces it uses.

That is why the default is the risk. Two controls are worth tightening now. Drop CAP_NET_RAW wherever the workload allows it, and treat Kubernetes checkpoint and restore permissions as privileged, because a checkpoint restored from an untrusted source can carry its old security context.

Volume makes this worse. DepthFirst counted 5,976 Linux kernel CVEs in 2026 to mid-September, with 1,650 in August alone. In Google’s kernelCTF, 13 of 36 disclosed bugs were reachable from unprivileged interfaces. The supply of escape routes is going up, not down.

The patch is not the remedy. A second kernel is.

Upstream fixed this on August 6. Ubuntu has not shipped it. Its security tracker listed 26.04 LTS as vulnerable and the fix in progress on September 21, and 24.04 LTS as vulnerable. The 22.04 LTS kernel is not affected. Ubuntu rates the issue Medium.

Updating container images does nothing here. The vulnerable code sits in the shared host kernel. The fix is a host kernel update and a reboot, or a live patch, on every node that runs a container.

The wider answer is to stop trusting the shared kernel for anything untrusted. DepthFirst points to a microVM, where each workload gets its own kernel. Firecracker and Kata Containers give every instance its own boundary, so a kernel escape stays inside one guest. We made this argument when Project Lunchbox put a kernel around each agent. It holds for a plain multi-tenant node too.

Three questions decide the exposure. Which node images still run an unpatched Ubuntu kernel. Who can call the checkpoint and restore API. And which workloads genuinely need to share a kernel with strangers.

Related reading. We covered the runtime layer where Kubernetes isolation keeps breaking, Edera’s per-agent kernel, and the quartet of Linux local privilege bugs that landed earlier in September.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.