Oct 6, 2026·5 min read·8 visits
A state desynchronization between vLLM's frontend and backend caches allows remote authenticated users to trigger a fatal assertion crash (CWE-617) by submitting a validation-failing multimodal request followed by a duplicate request, leading to complete server denial of service.
A state desynchronization (cache drift) vulnerability exists in the multimodal Inter-Process Communication (IPC) Least Recently Used (LRU) caches of vLLM. When a multimodal request fails validation after its media hash has been registered on the frontend but before the payload is committed to the backend engine core, the frontend and backend caches drift out of lockstep. A subsequent request reusing the same media triggers an assertion failure in the backend engine core, resulting in a complete denial of service.
vLLM uses a multi-process architecture to decouple request scheduling from high-performance GPU tensor computations. To optimize the processing of multimodal data such as images and video, vLLM implements a mirrored Inter-Process Communication (IPC) cache system.
This architecture consists of a frontend sender cache (MultiModalProcessorSenderCache on process P0) and an engine core receiver cache (MultiModalReceiverCache on process P1). The frontend cache acts as a shadow metadata tracker, while the backend cache stores the actual tensor payloads in GPU memory.
The system operates under a lockstep assumption. If the frontend registers a cache hit for a media hash, it transmits the request metadata with the payload data omitted (data=None), assuming the backend already possesses the cached tensor. A breakdown in this lockstep coordination exposes a reachable assertion bug class (CWE-617), allowing an attacker to terminate the backend engine process via network requests.
The root cause of the vulnerability is an unidirectional state mutation during the rendering phase that is not rolled back when downstream request validation fails.
When a multimodal request is processed, P0 renders the media, computes its unique cryptographic hash (mm_hash), and commits it directly to the P0 shadow metadata cache. If the request is subsequently rejected during validation (for example, if the prompt length exceeds the model's maximum limit and raises a ValueError), the request is discarded. Consequently, the actual tensor payload is never transmitted to or registered in P1's receiver cache.
This sequence causes state desynchronization: P0 assumes P1 has cached the media, whereas P1 has no record of the hash. When a subsequent request is submitted with the identical media, P0 detects a shadow cache hit and forwards the request with data=None. Upon receiving the request, P1 queries its receiver cache, returns None, and hits a fatal assertion block in vllm/multimodal/cache.py which terminates the process.
In vulnerable versions of vLLM (prior to 0.28.0), the engine core unconditionally assumed cache consistency. The receiver cache lookup in vllm/multimodal/cache.py contained the following check:
# Vulnerable assertion in vllm/multimodal/cache.py
if (cached_item := self._cache.get(mm_hash)) is not None:
return cached_item
assert mm_item is not None, f"Expected a cached item for {mm_hash=}"If the frontend registered a shadow cache hit, mm_item was received as None. If mm_hash was missing from self._cache due to previous request validation failures, the assertion failed, causing the engine to abort execution.
The vulnerability was resolved by implementing a dual mitigation strategy. First, the fatal assertion was replaced with a retryable exception in PR #46747:
# Patched implementation with recovery exception
if (cached_item := self._cache.get(mm_hash)) is not None:
return cached_item
if mm_item is None:
raise MultiModalCacheMissError([mm_hash])Second, immediate cache rollback was introduced in the OpenAI entrypoint (vllm/entrypoints/openai/chat_completion/serving.py) via PR #51897 to discard cache entries if a ValueError occurs during token validation:
# Patched OpenAI serving entrypoint
try:
max_tokens = get_max_tokens(...
)
except ValueError:
for rendered_input in engine_inputs:
if mm_hashes := cast(
MultiModalHashes | None, rendered_input.get("mm_hashes")
):
self.renderer.discard_mm_cache_entries(mm_hashes)
raiseAn attacker can systematically exploit this flaw through a two-step request sequence targeting any multimodal model served by a vulnerable vLLM instance.
First, the attacker submits a request containing an image payload paired with an excessively long prompt designed to trigger a validation failure (such as exceeding max_model_len). P0 computes the image hash, adds it to the shadow cache, and then the validation fails. The backend rejects the request, leaving its receiver cache empty.
Second, the attacker submits a valid request containing the exact same image. P0 identifies the hash in its shadow cache, sends data=None to the backend, and the backend crashes upon encountering the missing reference. This crash halts inference operations for all active clients on the shared server.
The execution of this vulnerability results in a complete denial of service of the vLLM inference server. Because vLLM is typically deployed as a shared daemon serving multiple downstream applications or clients, crashing the engine core process interrupts all active inference sessions.
This vulnerability has a CVSS 3.1 base score of 6.5 (Medium), reflecting the requirement for network access with low privileges (PR:L) and the resulting high availability impact (A:H). While confidentiality and integrity remain unaffected, the ease of triggering this state desynchronization makes it a highly effective vectors for targeted service disruptions.
Remediation requires upgrading the vLLM installation to version 0.28.0 or later, which integrates both the self-healing fallback protocol and proactive validation rollback.
If upgrading is not immediately possible, organizations can apply the following workarounds:
Disable multimodal IPC caching by adjusting the engine configuration parameters, which forces full payload delivery for every request and eliminates the lockstep dependency.
Implement upstream validation of prompt token lengths at the API gateway or proxy layer. This ensures that validation-failing requests are dropped before they reach vLLM's processing pipeline, preventing the initial cache desynchronization.
CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H| Product | Affected Versions | Fixed Version |
|---|---|---|
vllm vllm-project | < 0.28.0 | 0.28.0 |
| Attribute | Detail |
|---|---|
| CWE ID | CWE-617 (Reachable Assertion) |
| Attack Vector | Network (AV:N) |
| CVSS Score | 6.5 |
| Impact | Complete Denial of Service (DoS) |
| Exploit Status | Proof of Concept (PoC) / Unit Tests available |
| KEV Status | Not listed in CISA KEV |
The software contains an assertion statement that can be triggered by an input from an untrusted source, causing the application to terminate abnormally or crash, leading to a denial of service.
A critical remote code execution vulnerability (CVE-2026-102828) exists in simple-git versions 3.15.0 through 4.0.0. The vulnerability is caused by an incomplete blocklist within the library's default safety enforcement plugin, blockUnsafeOperationsPlugin. Attackers who can control Git configuration arguments or supply command flags to rebase operations can execute arbitrary system commands with the privileges of the parent Node.js process.
A critical security control bypass vulnerability exists in @simple-git/argv-parser before version 2.0.1. The package fails to map the VISUAL environment variable to the allowUnsafeEditor rule, allowing attackers who control environment parameters to execute arbitrary commands when Git triggers an interactive editor fallback.
A vulnerability in vLLM prior to 0.30.0 allows an authenticated multi-tenant attacker to infer execution history and prompt structures of other tenants. The multi-turn Responses API ('Harmony' path) fails to propagate the 'cache_salt' parameter during tool-call continuation steps, storing sensitive prompt prefixes in the global, unsalted cache space.
CVE-2026-105750 is a medium-severity local file disclosure vulnerability affecting the Docling and Docling-Slim libraries. When processing HTML documents using the optional Playwright rendering backend, the application fail to validate and restrict request URIs using the file:// scheme. This permits an attacker supplying a crafted HTML file to access, render, and exfiltrate local system files.
CVE-2026-102598 is a security bypass and Denial of Service (DoS) vulnerability in the Werkzeug WSGI web application library. In versions prior to 3.1.9, the library's safe_join function fails to sanitize Windows reserved device names containing an empty NTFS Alternate Data Stream (ADS) marker (such as NUL:). This allows remote, unauthenticated attackers to trigger indefinite thread-blocking operations on Windows hosts, resulting in application-wide resource exhaustion.
A vulnerability in @graphql-tools/executor-legacy-ws prior to version 1.1.35 hardcodes the TLS rejectUnauthorized setting to false for outgoing secure WebSocket (wss://) connections. This defect allows unauthenticated remote attackers to perform Adversary-in-the-Middle (MitM) attacks, capturing or tampering with sensitive connection payloads and subscription data.