JFrog has disclosed a critical, unpatched vulnerability in LMCache, open-source software used to speed up large language model servers such as vLLM. Tracked as CVE-2026-105192 and rated 9.8 out of 10, the flaw allows an unauthenticated attacker to run code on a cache server. It affects LMCache from version 0.3.9 through 0.5.5, as well as the 0.5.6 release candidates and the development branch. No fixed version is available, and LMCache has not published a security advisory.
The vulnerability exists in LMCache's multiprocess mode, where the cache runs as a standalone server that LLM workers reach over ZeroMQ. The socket has no authentication, and one type of message is unpacked with pickle, a Python format that can execute code during decoding. The server unpickles the data before checking the message type, so a crafted message can run the sender's code with the privileges of the LMCache process. On official container images, that process runs as root, according to JFrog.
By default, the multiprocess server listens only on localhost, so it is not reachable from other machines. It becomes exposed when an operator starts it with a routable address, as in multi-node deployments; LMCache's example Kubernetes deployment listens on every network interface. Until a patch ships, JFrog recommends not assigning the server a routable address and keeping its port on the local machine or a trusted cluster network. A firewall that limits access lowers the risk but does not remove it, because any host that can still connect can run code. Separately, a related denial-of-service flaw in vLLM, CVE-2026-105756, has already been fixed, while six additional LMCache security reports from a GitHub user remain unconfirmed and have no fix.