· via dev.to (home feed)
Unpatched 9.8-rated LMCache flaw lets unauthenticated ZeroMQ messages run code
JFrog disclosed CVE-2026-105192, a 9.8-rated unauthenticated RCE in LMCache's multiprocess KV cache server, with no fixed release available. Network isolation is the only containment.
JFrog disclosed CVE-2026-105192 on 7 October 2026, a critical remote code execution flaw rated 9.8 that affects the cache server in LMCache, a layer that stores LLM attention key-value caches outside the model process so long prompts and restarted engines do not pay for prefill twice. According to a write-up on dev.to, one unauthenticated message to that server is enough to execute code with the privileges of the LMCache process, and the project's official container images run that process as root. No fixed version exists.
How the flaw works
LMCache's multiprocess mode, which the documentation also calls distributed mode, opens a ZeroMQ ROUTER socket so worker processes can register and share KV cache blocks. The socket requires no authentication. Messages arrive encoded as msgpack, and one extension code is handed to a deserializer that invokes Python's pickle library while the server is still parsing the request's arguments, before the handler for that message type runs at all.
That ordering is the whole vulnerability. Pickle payloads execute during deserialization, so code runs before anything inspects what kind of message arrived. The JFrog record, as reported by dev.to, puts the default transport port at 5555, and a single ZeroMQ DEALER message to it suffices.
The affected range starts with 0.3.9, the release that introduced the multiprocess ZMQ transport and its pickle-backed extension in October 2025, and runs through 0.5.5, the current stable release. The 0.5.6 release candidates and the development branch carry the flaw as well.
The documented deployment is the exposed one
By default the cache server binds to loopback, which keeps it unreachable from other hosts. Exposure follows a single setting: operators who pass a routable address, exactly what multi-node deployments do so peers can share cached blocks, make the port reachable.
LMCache's own Kubernetes guidance walks into that configuration deliberately. It describes a DaemonSet running one cache server per node, shared by several vLLM pods, with hostNetwork enabled so model pods can find the server through the node's own address. Combined with a server listening on every interface, that leaves the cache port open to anything that can route to the node. The same page confirms that when the engine-driven path is loaded, KV transfers between server and workers use a pickle-based path by default.
The dev.to piece argues nobody made an obvious mistake here; the problem is sequence. A performance feature that began as an in-process buffer became a shared service, the service acquired a port, and the port arrived without authentication.
Why teams run it anyway
The project published its own justification for the extra process and port. In a benchmark using Qwen3-235B-A22B-Instruct-2507-FP8 across eight H100 GPUs with a multi-turn chat workload, mean time to first token fell from 3.98 seconds with in-process offloading to 0.29 seconds with multiprocess mode, and mean decoding speed rose from 9.81 to 37.47 tokens per second.
Sizing scales with those gains: the documentation recommends giving the L1 cache whatever host memory remains after the operating system and model server, with a documented example of 60GB and a benchmark run at 400GB. A pool that large on a node warrants an inventory entry before anyone debates authentication.
Containment while no patch exists
JFrog's guidance is network work with honest limits: keep the multiprocess server off routable addresses, holding its port on loopback or a trusted cluster network. A firewall restricting who can reach the port lowers the risk, but the CVE record is explicit that it does not remove it, because any host still able to open a connection can execute code.
Detection is also unresolved. Neither the CVE record nor JFrog's advisory offers an indicator of prior compromise, so log review starts from the socket itself: which addresses connected to the cache port, and when.
There is a version trap worth naming. A related flaw, CVE-2026-105756, affects vLLM and caused a malformed cache salt value to crash the engine on deployments using the LMCache multiprocess connector. That one is fixed in vLLM 0.30.0. As dev.to notes, upgrading the model server closes it and leaves the cache server flaw exactly where it was.
Why it matters
For anyone running inference infrastructure, this disclosure turns the KV cache into an inventory question: which cache servers exist, in which mode, what address each binds to, which networks can reach the port, and what else lives on those networks. Until a fixed release ships, placement is the entire defense against an unauthenticated, 9.8-rated code execution bug that starts with root privileges in the official images. Teams serving models with vLLM should check their nodes today, because the documented deployment path is the exposed one.
- #security
- #llm
- #vllm
- #kubernetes
- #vulnerability