Executive Summary
A flaw in LMCache, the cache layer that makes vLLM inference faster, lets an attacker run commands on the cache host with no login. JFrog published it on October 7 as CVE-2026-105192 and scored it 9.8. The multiprocess mode that lets worker processes share cache blocks opens a ZeroMQ socket with no authentication, and the server passes a message from that socket straight into pickle.loads. One network message executes code as the user that runs the cache. The official container images run that process as root.
The exposure depends on how the deployment is configured. A stock single-host install binds to localhost and is not reachable from another machine. Any deployment that sets the host flag to a routable address so peers on other nodes can share cache is open to the network. There is no fixed release. Version 0.5.5, the 0.5.6 release candidates, and the development branch all still call pickle.loads on unauthenticated input. Until the transport is authenticated, the fix is architectural, and it is to keep the port on localhost.
An inference server spends most of its wall time recomputing prompt prefixes it has already seen. LMCache exists to stop that waste. It stores the key-value cache, the intermediate state a model builds while reading a prompt, so a later request that shares a prefix skips the work. On a busy deployment the saving is real, which is why the project now turns up across the serving stacks our readers run.
The project runs in two modes. In the simple mode it lives inside the vLLM process. In multiprocess mode, which the project also calls distributed mode, it runs as a standalone server and worker processes register cache blocks with it over ZeroMQ. The second mode is what multi-node serving uses. It is also where the bug lives.
A socket with no lock on it
The server opens a ZeroMQ ROUTER socket so worker processes can register and share KV cache blocks. The intended clients are sibling LMCache processes. The socket carries no CURVE, no ZAP, no password, and no message authentication. By default it binds to localhost. Set the host flag to a routable address, which multi-node deployments do, and the port answers anything on the network.
Messages on that socket are msgpack. Extension code 1 is registered for DeviceIPCWrapper, and the decode path sends that payload into DeviceIPCWrapper.Deserialize, which calls pickle.loads. That call runs while the server is still decoding the REGISTER_KV_CACHE arguments, before the handler that would do anything useful with them. Pickle runs arbitrary code by design, so one unauthenticated message to the transport port executes commands as the LMCache user. The advisory proof of concept is a single line, uid=0(root).

The gap shipped in 2025 and is still there
The advisory is blunt about the age of the mistake. The decode path shipped in version 0.3.9, in October 2025, and it remains in the latest public release, 0.5.5, in the 0.5.6 release candidates, and on the development branch as of October 7. No fixed version has been published. That is the same class of error researchers grouped as ShadowMQ in November 2025, where AI inference frameworks handed bytes from an unauthenticated socket to a deserializer. It is why the remedy here is a code change and not a patch.
The recommendation is the one the project should have shipped at the start. Stop passing network bytes to pickle. Replace the serializer behind the msgpack extension with a safe format, and authenticate the ZeroMQ transport with CURVE or a per-message HMAC. The advisory also suggests refusing to bind a routable address unless that authentication is configured, which is the right default for a cache server.
What to do this week, since patching is not an option
Skip the usual upgrade instruction. There is nothing to upgrade to. The mitigation is configuration. Do not set the host flag to a routable address. Keep the multiprocess port on localhost, or on a cluster network you already treat as trusted. A firewall narrows who can reach the port, but every host that can open a connection to it can still run code as the LMCache user. That is not a comfortable place to leave a service that sits next to your model weights.
Confirm what your stack actually runs. A cache process that only lives inside the vLLM process never opens the port. A standalone cache server that shares blocks across nodes does. The difference is one flag, and it is the difference between a local utility and a network service with no lock.
Ask three questions of your own environment. Does your inference stack run LMCache in multiprocess mode, or only inside the vLLM process? If it runs standalone, is the transport port bound to localhost or to a routable address? And what user does the cache process run as, because root turns a cache bug into a host takeover?
Related reading. Our October 8 piece on the NVIDIA DCGM exporter flaw covers a different service that is exposed by default on GPU nodes, and why a metrics endpoint belongs in the same threat model as the accelerator.
Get the next one before it is old news
Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.

[…] reading. The same lesson showed up in the cache server that feeds vLLM, where an unauthenticated flaw handed out a root shell, and in the exposed GPU metrics endpoints […]