Oct 6, 2026·7 min read·8 visits
vLLM fails to propagate the cache partition salt during multi-turn tool execution continuations, exposing post-tool prompt prefixes in the shared global cache and creating a side-channel oracle for co-tenants.
A vulnerability in vLLM prior to 0.30.0 allows an authenticated multi-tenant attacker to infer execution history and prompt structures of other tenants. The multi-turn Responses API ('Harmony' path) fails to propagate the 'cache_salt' parameter during tool-call continuation steps, storing sensitive prompt prefixes in the global, unsalted cache space.
vLLM is an open-source, high-throughput LLM serving engine designed to optimize inference performance using advanced memory management strategies. A core optimization within this engine is prefix caching, which retains the Key-Value (KV) cache of prompt prefixes across sequential inference requests. By caching the KV states of common prompts or system instructions, the engine significantly reduces the Time-to-First-Token (TTFT) and minimizes redundant graphics processing unit (GPU) computations during high-concurrency operations.
In multi-tenant environments where a single vLLM instance is shared among independent client accounts, sharing a global prefix cache exposes a side-channel vector. A malicious tenant can exploit timing patterns and metadata parameters to determine if other tenants have executed specific instructions. To prevent this cross-tenant information leakage, vLLM introduces the cache_salt configuration argument. The salt acts as a namespace partition by blending the salt value with the prompt prefix hash, thereby isolating cached KV states between different tenants.
CVE-2026-105752 (co-tracked under GitHub Advisory ID GHSA-935w-9g4m-p28p) identifies a flaw in this isolation architecture within vLLM versions prior to 0.30.0. Specifically, the vulnerability resides in the multi-turn GPT-OSS 'Harmony' path executed through the /v1/responses endpoint. When processing multi-turn interactions involving external tools or function calls, the engine fails to preserve and propagate the cache_salt parameter during automatic continuation turns. As a result, subsequent prompt states are stored in the global, unsalted prefix cache, defeating the privacy guarantees of the namespace partitioning mechanism.
The core of the vulnerability lies in how the vLLM multi-turn Responses API manages context transitions during tool execution. When a model request involves a tool or function call, the communication operates in an execution loop. The client initiates the first turn of the conversation, which correctly incorporates the client-specific cache_salt parameter, ensuring that the initial prompt is cached within an isolated partition.
Once the model triggers a tool, the engine processes the tool call locally, collects the output, and automatically initiates a continuation turn to feed the tool execution results back to the model context. During this transition, the handler must reconstruct the input objects for the next inference phase. However, the logic responsible for creating subsequent input structures fails to carry forward the original cache_salt argument from the initial query context.
Because the cache_salt is omitted during the reconstruction of the continuation engine inputs, the engine defaults to the global, unsalted namespace. Consequently, the prefix representing the post-tool execution state is cached in the shared storage area rather than the tenant-isolated sandbox. This omission compromises the cryptographic isolation of the cache and creates a reliable side-channel oracle for co-tenants.
An analysis of the vulnerable codebase in vllm/entrypoints/openai/responses/serving.py confirms that the continuation logic was completely blind to the cache_salt configuration. During the evaluation loop of a multi-turn conversation, the context was updated via the HarmonyContext rendering functions, but the instantiation of tokens_input occurred without parameter forwarding:
# PRE-FIX VULNERABLE CODE
# Location: vllm/entrypoints/openai/responses/serving.py
if isinstance(context, HarmonyContext):
token_ids = context.render_for_completion()
# BUG: The tokens_input helper is called without passing the cache_salt
engine_input = tokens_input(token_ids)
sampling_params.max_tokens = max_model_len - len(token_ids)The patch implemented in commit 6a2a2bb02b563b83f946012959fd3927984d072a fixes the issue by extracting the cache_salt value from the original request inputs and passing it down to a newly developed rendering routine on the online_renderer object:
# POST-FIX PATCHED CODE
# Location: vllm/entrypoints/openai/responses/serving.py
# The fix extracts the salt from the first turn of the engine input
cache_salt = cast(str | None, engine_input.get("cache_salt"))
# ...
if isinstance(context, HarmonyContext):
# The cache_salt is explicitly passed to the online renderer
engine_input = self.online_renderer.render_responses_harmony_messages(
context.messages,
cache_salt=cache_salt,
tok_params=tok_params,
)
sampling_params.max_tokens = max_model_len - self._extract_prompt_len(
engine_input
)Additionally, vllm/renderers/online_renderer.py was extended to include the render_responses_harmony_messages helper method, which binds the extracted cache_salt directly into the final tokens_input parameters during downstream processing. This guarantees that all subsequent turns are correctly partitioned inside the isolated namespace:
# POST-FIX PATCHED RENDERER
# Location: vllm/renderers/online_renderer.py
def render_responses_harmony_messages(
self,
messages: list[OpenAIMessage],
*,
cache_salt: str | None,
tok_params: TokenizeParams | None = None,
) -> EngineInput:
arrival_time = time.time()
prompt = TokensPrompt(prompt_token_ids=render_for_completion(messages))
if tok_params is not None:
tok_params.apply_post_tokenization(
self.renderer.tokenizer,
prompt,
)
# The salt is applied to the output engine_input structure
engine_input = tokens_input(prompt["prompt_token_ids"], cache_salt=cache_salt)
engine_input["arrival_time"] = arrival_time
return engine_inputTo exploit this vulnerability, an attacker must have low-privileged API access on the same multi-tenant vLLM server as the targeted victim. The attacker must also possess sufficient understanding of the victim's application workflow to reconstruct or predict the low-entropy structure of the post-tool execution prompt. This scenario is highly plausible in systems running standardized agent templates or commercial frameworks.
First, the victim issues a multi-turn request that incorporates a specific cache_salt and invokes a tool call. The vLLM engine processes the tool, drops the salt during the continuation turn, and caches the resulting post-tool prefix in the global, unsalted prefix cache namespace. Because the metadata contains standardized formatting around tool outputs, the cached prefix becomes predictable.
Next, the attacker constructs a probe request using the anticipated post-tool prompt structure and submits it to the /v1/responses endpoint. By analyzing the returned response, the attacker inspects metadata attributes such as cached_tokens_per_turn or evaluates the response latency. A non-zero cache count or significantly faster response time indicates a cache hit, verifying that another tenant executed that specific tool workflow. This timing side-channel acts as a precise membership oracle.
The security impact of CVE-2026-105752 is classified as a low-severity information leak. Although the vulnerability does not allow attackers to directly exfiltrate model weights, raw database records, or explicit conversation transcripts, it breaks the core tenant isolation boundaries defined by the prefix cache architecture.
The primary danger is metadata leakage and behavioral monitoring. An attacker can determine whether other users are interacting with sensitive internal tools, executing specific database queries, or generating routine financial transactions. This metadata exposure is highly critical in enterprise multi-tenant deployments where strict operational privacy is required.
The CVSS v3.1 score of 3.1 reflects the operational complexity of the attack. Exploitation requires high attack complexity because the attacker must anticipate the prompt formatting and have active low-privilege co-tenancy access. However, because prefix caching is enabled by default to optimize infrastructure costs, the affected attack surface is broad for setups leveraging the Harmony Responses framework.
The definitive fix for this vulnerability is to upgrade the vLLM installation to version 0.30.0 or later. The update introduces the required code alterations to ensure that the cache_salt is persistently propagated through all internal continuation loops of the multi-turn processing lifecycles.
If upgrading is not immediately possible in production environments, administrators can mitigate the issue by disabling prefix caching globally. To apply this change, start the vLLM engine without the caching flag or set the parameter explicitly: --enable-prefix-caching=false. Note that disabling this parameter will degrade latency (specifically Time-to-First-Token) and increase computational costs under high workloads.
Additionally, security teams should implement monitoring controls on the vLLM server. Track the distribution of prompt cache hits across distinct tenant API tokens. An anomalous density of cache hits matching standard agent templates across different user accounts may indicate active probing or membership enumeration attempts.
CVSS:3.1/AV:N/AC:H/PR:L/UI:N/S:U/C:N/I:L/A:N| Product | Affected Versions | Fixed Version |
|---|---|---|
vLLM vllm-project | < 0.30.0 | 0.30.0 |
| Attribute | Detail |
|---|---|
| CWE ID | CWE-524, CWE-200 |
| Attack Vector | Network |
| CVSS v3.1 Score | 3.1 (Low) |
| EPSS Score | 0.00043 |
| Impact | Partial Tenant Cache Isolation Failure / Metadata Leakage |
| Exploit Status | none |
| KEV Status | Not Listed |
The product uses a cache that contains sensitive information, but the cache is not properly protected or isolated, allowing unauthorized actors to access or infer the cached data.
A critical remote code execution vulnerability (CVE-2026-102828) exists in simple-git versions 3.15.0 through 4.0.0. The vulnerability is caused by an incomplete blocklist within the library's default safety enforcement plugin, blockUnsafeOperationsPlugin. Attackers who can control Git configuration arguments or supply command flags to rebase operations can execute arbitrary system commands with the privileges of the parent Node.js process.
A critical security control bypass vulnerability exists in @simple-git/argv-parser before version 2.0.1. The package fails to map the VISUAL environment variable to the allowUnsafeEditor rule, allowing attackers who control environment parameters to execute arbitrary commands when Git triggers an interactive editor fallback.
A state desynchronization (cache drift) vulnerability exists in the multimodal Inter-Process Communication (IPC) Least Recently Used (LRU) caches of vLLM. When a multimodal request fails validation after its media hash has been registered on the frontend but before the payload is committed to the backend engine core, the frontend and backend caches drift out of lockstep. A subsequent request reusing the same media triggers an assertion failure in the backend engine core, resulting in a complete denial of service.
CVE-2026-105750 is a medium-severity local file disclosure vulnerability affecting the Docling and Docling-Slim libraries. When processing HTML documents using the optional Playwright rendering backend, the application fail to validate and restrict request URIs using the file:// scheme. This permits an attacker supplying a crafted HTML file to access, render, and exfiltrate local system files.
CVE-2026-102598 is a security bypass and Denial of Service (DoS) vulnerability in the Werkzeug WSGI web application library. In versions prior to 3.1.9, the library's safe_join function fails to sanitize Windows reserved device names containing an empty NTFS Alternate Data Stream (ADS) marker (such as NUL:). This allows remote, unauthenticated attackers to trigger indefinite thread-blocking operations on Windows hosts, resulting in application-wide resource exhaustion.
A vulnerability in @graphql-tools/executor-legacy-ws prior to version 1.1.35 hardcodes the TLS rejectUnauthorized setting to false for outgoing secure WebSocket (wss://) connections. This defect allows unauthenticated remote attackers to perform Adversary-in-the-Middle (MitM) attacks, capturing or tampering with sensitive connection payloads and subscription data.