Critical LMCache Vulnerability Exposes AI Servers to RCE Attacks

LMCache vulnerability CVE-2026-105192 ZeroMQ Python pickle RCE

A severe vulnerability recently emerged within LMCache, a sophisticated acceleration system designed for large language models. This profound flaw permits malicious actors to execute arbitrary commands on a server without requiring a password. Under specific configurations, an assailant can astonishingly acquire full root privileges. Alarmingly, an official corrected version remains unavailable at this time.

The Discovery of CVE-2026-105192

Yuval Moravchick, an esteemed researcher from JFrog Security Research, originally uncovered this perilous problem. Consequently, the vulnerability received the official identifier CVE-2026-105192, alongside a critical severity score of 9.8 out of 10 on the CVSS 3.1 scale. Frighteningly, this attack demands no user account, special privileges, or user interaction whatsoever. Merely possessing network access to the vulnerable service, when operating in distributed mode, proves entirely sufficient.

LMCache significantly accelerates neural network response generation by meticulously preserving the KV-cache. This cache contains intermediate computational results that the language model can efficiently reuse later. Thanks to this ingenious mechanism, the server effectively avoids reprocessing familiar textual fragments. Developers frequently employ this technology alongside massive language model deployment systems, such as vLLM. It proves particularly invaluable when processing lengthy queries and serving multitudes of concurrent users.

The Danger of Distributed Mode and ZeroMQ

The dangerous flaw resides specifically within the LMCache distributed mode. In this mode, individual worker processes rapidly exchange cache contents utilizing ZeroMQ. The server opens a specialized ROUTER network socket, typically designated on port 5555, to receive messages from other system components. However, the developers inexplicably failed to implement mandatory authentication for this crucial connection. Thus, any node capable of connecting to the port can freely transmit messages without its identity ever being verified.

The Peril of Insecure Deserialization

The primary issue stems directly from insecure deserialization within the LMCache system. This occurs when a program reconstructs complex objects from received data streams. Although LMCache utilizes the msgpack format, one specific message type tragically passes contents directly to the Python language’s pickle.loads function. Unlike standard data reading procedures, the pickle mechanism possesses the terrifying capability to execute arbitrary code while reconstructing objects. Consequently, a meticulously crafted packet instantaneously transforms into a lethal command for the host server’s operating system.

The vulnerable segment lies deeply embedded within the DeviceIPCWrapper.Deserialize handler. When the server parses the arguments of a REGISTER_KV_CACHE message, the msgpack extension carrying code 1 invokes pickle.loads. Disastrously, this occurs before the primary handler ever executes or verifies the received values. Therefore, a malicious actor merely needs to dispatch a single, carefully prepared ZeroMQ DEALER message. Even a subsequent processing error cannot prevent the execution of the embedded malicious command.

Achieving Root Access in Container Environments

JFrog conclusively verified the exploitability of this flaw by deploying a functional attack prototype. The astute researchers successfully forced the server to execute a system command and meticulously record the result within a file. Their verification shockingly revealed that the compromised process utilized the omnipotent root account, possessing a UID of 0. The reason for this heightened danger lies directly within the official LMCache container images, where the corresponding process stubbornly launches with maximum privileges.

Mitigation Strategies and Current Status

However, this critical CVSS score does not imply that every LMCache installation remains susceptible to remote hacking. By default, the ZeroMQ transport strictly listens only to localhost. Therefore, external computers cannot connect to the service directly. A remote attack only becomes viable when an administrator explicitly assigns a network-accessible address utilizing the --host parameter. Administrators typically employ this configuration when deploying LMCache across multiple servers. If the component operates exclusively within a single vLLM process, the vulnerable network port remains safely closed.

This problematic mechanism originally appeared in LMCache version 0.3.9 and stubbornly persisted throughout subsequent releases. At the precise moment of the vulnerability’s disclosure on October 7, 2026, the latest stable iteration remained version 0.5.5, which was published on September 12. The flaw also haunted release candidates up to version 0.5.6rc3, and it lingered within the development branch as of October 7. Crucially, as of October 8, no published patch could be found.

Researchers urgently advise administrators to immediately audit their distributed LMCache server configurations and drastically restrict access to the transport port. Until a comprehensive patch emerges, the safest approach involves leaving ZeroMQ accessible exclusively via localhost or from a strictly controlled internal network. While a robust firewall effectively reduces the pool of potential attackers, it ultimately fails to eradicate the vulnerability itself. Any network participant possessing authorized access to the port retains the terrifying capability to transmit a dangerous message.

Support Our Threat Intelligence

If you find our technology report and cybersecurity news helpful, consider supporting our work.

Crypto QR Code
USDT (TRC20):
TN8BdV8cp4T1Cd28gK9qTAnZknzzuwyUtm
USDT (ERC20):
0x3725e1a7d3bc5765499fa6aaafe307fabcd75bce

Leave a Reply