Oct 8, 2026·6 min read·6 visits
Docling document parsers lack validation checks for table row and column spans, allowing tiny documents with extreme span values to exhaust host memory and CPU resources.
An uncontrolled resource consumption vulnerability exists in the Docling document conversion library. Maliciously structured HTML, JATS, ODS, or BoxNote inputs containing table cells with excessively large 'rowspan' or 'colspan' attribute values trigger algorithmic complexity conditions. This allows unauthenticated remote attackers to initiate resource exhaustion states, crashing or hanging the target document processing pipeline while bypassing configured timeouts.
Docling is an enterprise-oriented, open-source Python library used to parse, segment, and convert complex document formats into structured, machine-readable formats. It serves as a critical component in automated ingestion pipelines, Retrieval-Augmented Generation (RAG) frameworks, and Large Language Model (LLM) workflows.
The library supports various data backends, adapting its conversion mechanisms for different input file types. These backends parse structured tables (such as HTML, JATS XML, OpenDocument Spreadsheet, and BoxNote JSON documents) to build structured layout matrices. To properly model these tables, the parsers reconstruct the physical arrangement of overlapping grid structures.
In vulnerable versions (prior to 2.131.0), the parsing adapters accepted row and column span attributes without verifying their scale against reasonable upper bounds. This missing validation exposes applications using Docling to uncontrolled resource exhaustion attacks. When processing a crafted file with highly inflated span coordinates, the target system encounters complete resource starvation.
The vulnerability stems from the combined effects of Uncontrolled Resource Consumption (CWE-400) and Memory Allocation with an Excessive Size Value (CWE-789). When Docling ingests tables, the underlying backend adapter instantiates a two-dimensional grid array to map spans and calculate intersecting coordinates. This mapping ensures subsequent processing steps maintain precise cell layout tracking.
Prior to the patch, when the parser processed attributes such as 'rowspan' or 'colspan', it converted the raw metadata strings directly into integers. The parser utilized these values to drive nested execution loops that initialized virtual table cells and populated spatial coordinate sets. An attacker can set these integers to extreme values, forcing the execution of intensive loops that map millions of coordinate intersections.
Because Python objects carry measurable heap memory overhead, allocating large tables or tracking their indexes in coordinate loops leads to quick memory exhaustion. The execution time and memory footprint scale exponentially relative to the magnitude of the supplied span attributes.
Furthermore, Docling provides a 'document_timeout' feature to limit execution runs, but this mechanism is only evaluated between pipeline phases. Because the synchronous table-parsing logic executes inside a single continuous phase, the parser cannot yield control to the scheduler. The timeout check is bypassed entirely, leaving the system to run until it is terminated by an out-of-memory handler or host system starvation.
The remediation implemented in commit c5b4429cc6500a344c13edeb22e67610c2159b09 addresses the root cause of the resource exhaustion by introducing a static-limit check and context-aware dynamic clamping.
The developers created docling/backend/utils/table_spans.py to define static limits for table attributes:
# Standard limits introduced to prevent algorithmic complexity attacks
MAX_COLSPAN = 1000
MAX_ROWSPAN = 65534
def clamp_span(span: int, limit: int) -> int:
"""Clamp a declared span to the safe range [1, limit]."""
return max(1, min(span, limit))To prevent a secondary CPU consumption vector—where an exceptionally long digit string causes resource exhaustion during string-to-integer conversion—the code restricts the length of processed numeric substrings before conversion. This logic is implemented in the _extract_num method in html_backend.py:
# Pre-validation check of string lengths to prevent parsing exhaustion
digits = match.group().lstrip('0')
if len(digits) > len(str(MAX_ROWSPAN)):
return MAX_ROWSPAN
return max(int(digits or '0'), 1)In addition to static limits, the backends apply dynamic clamping based on the table's physically detected layout. This context-aware bounding is implemented across the distinct parser adapters as shown below:
# Dynamic layout clamping inside boxnote_backend.py
row_span = min(
_declared_span(cell, 'rowspan', MAX_ROWSPAN),
len(rows) - row_idx,
)
col_span = min(
_declared_span(cell, 'colspan', MAX_COLSPAN),
max(width - col_idx, 1),
)By ensuring that cells are constrained to both safe static boundaries and the physical boundaries of the table region, the processing complexity remains directly proportional to the actual structural properties of the document.
An attack targeting this vulnerability requires zero authentication and minimal operational overhead. The vulnerability can be triggered remotely by supplying a crafted document file to any application utilizing the Docling library for document parsing or automated content ingestion.
A malicious actor can construct a small HTML document containing highly disproportionate table properties to target the underlying HTMLDocumentBackend:
<table>
<tr>
<td rowspan='999999999'>Payload Column A</td>
<td colspan='999999999'>Payload Column B</td>
</tr>
<tr>
<td>Additional Row</td>
</tr>
</table>When Docling ingests this file, the conversion process attempts to initialize and iterate over a matrix of size 999999999 x 999999999. This operation triggers immediate heap-allocation calls and locks the active thread in a continuous execution loop.
The exploitation characteristics are highly asymmetric. The malicious payload requires less than a kilobyte of network bandwidth to transmit, yet it causes gigabytes of memory allocation on the server side within milliseconds, leading to a denial of service.
The successful exploitation of CVE-2026-105749 results in a complete denial of service for the host processing system. Because the vulnerability targets fundamental document parsers, its downstream impact is amplified in automated production environments.
In Retrieval-Augmented Generation (RAG) environments, document parsing pipelines often ingest and chunk large volumes of user-supplied business files. A single malicious upload can crash the parsing service, disabling the ingestion pipeline for all legitimate users. This risk is especially critical in multitenant applications where document ingestion shares physical host resources with core application components.
Because the host process memory limit is quickly exceeded, the operating system's Out-of-Memory (OOM) killer will typically terminate the parent process. In microservice architectures, such as Kubernetes clusters, continuous container crashes can trigger CrashLoopBackOff states. This destabilizes the local service cluster, resulting in broader infrastructure degradation.
This vulnerability has been assigned a CVSS score of 6.5 (Medium), reflecting high availability impact but low confidentiality or integrity exposure. Because exploiting the flaw requires zero authentication or specific privileges, it is highly attractive to adversaries seeking to execute low-cost denial of service attacks.
Remediation requires upgrading the docling package to version 2.131.0 or higher, which imports docling-core version 2.98.0 or higher. Organizations should verify their current python environment version using pip dependency tools:
pip list | grep doclingIf upgrading cannot be implemented immediately, systems can deploy input validation filters to analyze incoming files before they are processed by Docling. Organizations should screen uploaded XML, HTML, and ODS content using regular expressions to detect anomalous span configurations:
(colspan|rowspan)\s*=\s*['"](9{4,}|\d{5,})['"]Additionally, parsing operations should be isolated within containers with strict memory limits enforced by control groups (cgroups). This deployment pattern ensures that an Out-of-Memory event terminates only the worker container, protecting the parent system and neighboring processes from cascading failures.
CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:U/C:N/I:N/A:H| Product | Affected Versions | Fixed Version |
|---|---|---|
docling IBM / docling-project | >= 2.0.0, < 2.131.0 | 2.131.0 |
docling-slim IBM / docling-project | >= 2.92.0, < 2.131.0 | 2.131.0 |
| Attribute | Detail |
|---|---|
| CWE ID | CWE-400 (Uncontrolled Resource Consumption) |
| Attack Vector | Network / Remote Document Ingestion |
| CVSS Severity | 6.5 (Medium) |
| EPSS Score | 0.00256 (Percentile: 15.73%) |
| Impact | Complete Application Hang or Out-of-Memory Crash |
| Exploit Status | Proof-of-Concept Available |
| KEV Status | Not Listed |
The software does not properly control the allocation and maintenance of a limited resource, enabling an actor to cause a denial of service.
Docling prior to version 2.131.0 is vulnerable to arbitrary local code execution during module initialization due to incorrect order of operations in its plugin discovery system. Even when the default option to reject external plugins is active, Docling utilizes Pluggy to scan and import entrypoints before performing namespace validation.
A Local File Inclusion (LFI) and Arbitrary File Disclosure vulnerability exists in Docling and Docling Slim versions >= 2.16.0 up to 2.131.0. When parsing serialized DoclingDocument structures using the JSON input format, the backend fails to restrict image URI schemes, allowing remote attackers to retrieve local files and verify path existence on the host system during embedded document export.
Docling, a tool for parsing and processing diverse document formats, is vulnerable to arbitrary file read, arbitrary file write, and potential remote code execution (RCE) in versions 2.94.0 through 2.131.0. The vulnerability occurs when applications configure Docling to use the Tectonic engine for rendering TikZ diagrams into images. Because the compilation did not restrict hazardous TeX primitives or sandbox the environment, an attacker can supply crafted documents containing malicious TikZ definitions to access or modify local files and execute arbitrary commands under the privileges of the processing application.
An SSRF guard bypass vulnerability in the Docling document conversion engine allows unauthenticated attackers to bypass internal IP access controls. The vulnerability exists due to a DNS rebinding Time-of-Check Time-of-Use (TOCTOU) condition, URL authority parsing inconsistencies, and unvalidated network requests triggered during headless browser page rendering.
A technical analysis of CVE-2026-105742 (GHSA-p3fw-7699-7926), a sensitive information disclosure vulnerability in the Docling document processing library. Vulnerable versions of Docling indiscriminately forward custom HTTP headers, such as authentication tokens, to arbitrary third-party origins and during cross-origin redirects while fetching remote image assets from untrusted HTML and EPUB documents.
CVE-2026-106121 is a Denial of Service (DoS) vulnerability in the RabbitMQ Java Client library (amqp-client) affecting versions prior to 5.37.0. The vulnerability resides in the legacy, custom JSON-RPC parsing class com.rabbitmq.tools.json.JSONReader. When parsing malformed or truncated payloads ending within a quoted string or single-line comment, the parser's scanner enters an infinite loop. This occurs because the loop lacks an exit condition for the end-of-input sentinel character returned by the iterator, leading to either CPU exhaustion or a JVM crash from an OutOfMemoryError.