Aug 7, 2026·7 min read·61 visits
A vulnerability in pypdf allows attackers to trigger a Denial of Service (OOM crash) via crafted /ToUnicode CMap streams containing massive hex-encoded tokens.
An uncontrolled resource consumption vulnerability (CWE-400) exists in pypdf prior to version 6.15.0. When extracting text from a specially crafted PDF document, the parser fails to restrict token lengths within /ToUnicode CMap streams, causing unbounded memory allocation and process termination via Out-of-Memory (OOM) crashes.
The pypdf library is a widely deployed pure-Python library utilized for parsing, manipulating, and extracting text from PDF documents. A core feature of this library is text extraction, which translates internal PDF character codes into standardized Unicode glyphs. This translation relies heavily on decoding /ToUnicode character map (CMap) streams embedded within font descriptors. These streams define how custom or non-standard encodings map to the Unicode coordinate space.
The vulnerability, designated as CVE-2026-71870, is classified under CWE-400 (Uncontrolled Resource Consumption). The vulnerability is localized in the CMap parsing routines of the library, specifically within the handling of space-separated tokens in bfrange definitions. When processing untrusted files, the library fails to validate or restrict the byte length of tokens parsed from these streams.
This lack of restriction exposes any service running pypdf to resource exhaustion. Exposed applications typically include automated document indexing platforms, email attachment scanners, search engine crawlers, and optical character recognition (OCR) preprocessors. An attacker can craft a document that, when indexed or parsed, consumes all available system memory, leading to a process-level Out-of-Memory (OOM) event and triggering a denial-of-service condition.
The root cause of CVE-2026-71870 lies within the parsing loop of /ToUnicode CMap streams, specifically inside the parse_bfrange function located in pypdf/_cmap.py. A standard CMap definition maps ranges of character codes to destinations using blocks defined by the beginbfrange and endbfrange operators. The parser reads these blocks sequentially and splits entries into tokens using space character delimiters.
In vulnerable versions of pypdf (prior to 6.15.0), the parser processes these string tokens dynamically using operations like unhexlify and string interpolation without validating token length. An attacker can inject an arbitrarily long hexadecimal string enclosed in angle brackets (such as <4141...41>) within a bfrange block. Because the library performs multi-line parsing and processes every element of the split line, it attempts to load, decode, and map these extremely long sequences.
During processing, pypdf attempts to decode the token using the utf-16-be or charmap codecs and stores the mapping in an internal dictionary (map_dict). Python dictionaries incur memory overhead for each key-value pair, and handling multi-megabyte string objects during runtime dramatically increases heap size. A crafted document under 100 KB in size can trigger allocation pools in the gigabyte range, causing rapid exhaustion of the virtual memory space.
The vulnerability is localized within pypdf/_cmap.py. In affected versions, the parsing implementation for parse_bfrange extracts parameters from lines directly, splitting on byte-level spaces. Below is the vulnerable parsing code path:
# Affected code path in pypdf/_cmap.py prior to 6.15.0
lst = [x for x in line.split(b" ") if x]
a = int(lst[0], 16)
b = int(lst[1], 16)
...
# Unbounded byte manipulation and mapping assignment
map_dict[unicode_key] = unhexlify(sq).decode("utf-16-be", "surrogatepass")The patch implemented in commit afba8080e19d29a3c256a742b340995e695b35aa addresses the lack of bounds enforcement by introducing a validation utility and strictly defined byte limits. Because hexadecimal representation doubles the required characters, the token limits are defined as double the maximum byte count:
# Patched limits and validation helper in pypdf/_cmap.py
MAX_CMAP_CODE_BYTES = 8
MAX_CMAP_STRING_BYTES = 512
MAX_CMAP_CODE_BYTES_LIMIT = MAX_CMAP_CODE_BYTES * 2
MAX_CMAP_STRING_BYTES_LIMIT = MAX_CMAP_STRING_BYTES * 2
def _check_token_length(token: bytes, limit: int) -> None:
token_length = len(token)
if token_length > limit:
description = {
MAX_CMAP_CODE_BYTES_LIMIT: "code",
MAX_CMAP_STRING_BYTES_LIMIT: "string",
}.get(limit, "token")
raise LimitReachedError(
f"Maximum /ToUnicode {description} length exceeded: {token_length} > {limit}."
)The validation checks are embedded into each critical stage of the parse_bfrange loop, assessing code-length parameters (checked against MAX_CMAP_CODE_BYTES_LIMIT) and multi-byte destination strings (checked against MAX_CMAP_STRING_BYTES_LIMIT):
# Code-level insertion of token validation inside parse_bfrange
if multiline_rg is not None:
for sq in lst[3:]:
if sq == b"]":
break
_check_token_length(sq, limit=MAX_CMAP_STRING_BYTES_LIMIT)
else:
_check_token_length(lst[0], limit=MAX_CMAP_CODE_BYTES_LIMIT)
_check_token_length(lst[1], limit=MAX_CMAP_CODE_BYTES_LIMIT)
if lst[2] == b"[":
for sq in lst[3:]:
if sq == b"]":
break
_check_token_length(sq, limit=MAX_CMAP_STRING_BYTES_LIMIT)
else:
_check_token_length(lst[2], limit=MAX_CMAP_STRING_BYTES_LIMIT)To exploit this vulnerability, an attacker constructs a PDF document containing a font descriptor object with a crafted /ToUnicode mapping. The stream within this mapping contains a bfrange block that defines mappings using abnormally long sequences of hexadecimal characters inside the < and > delimiters.
An example configuration of an payload block within the CMap stream is structured as follows:
/CIDInit /ProcSet findresource begin
12 dict begin
begincmap
/CIDSystemInfo << /Registry (Adobe) /Ordering (UCS) /Supplement 0 >> def
/AMapName /Unicode def
1 begincodespacerange
<0000> <FFFF>
endcodespacerange
1 beginbfrange
<0000> <0001> <4141414141414141...[repeated 100000 times]...4141>
endbfrange
endcmap
CMapName currentdict /CMap defineresource pop
end
endThe target application processes this stream during standard operations that invoke text extraction. For example, if a developer writes an automation script to index uploaded resumes or documents, the execution of the following Python block triggers the vulnerability:
from pypdf import PdfReader
# Opening and parsing the malicious file structure
reader = PdfReader("malicious.pdf")
for page in reader.pages:
# This call parses the ToUnicode stream and exhausts memory
extracted_text = page.extract_text() Because the parsing occurs synchronously within the application thread, the interpreter process dynamically inflates in heap size until the kernel's Out-Of-Memory (OOM) killer or the runtime manager terminates the process. No authentication or privileged session is required to trigger this crash.
The impact of CVE-2026-71870 is confined to service availability. Because this is an uncontrolled resource consumption issue (CWE-400), it does not directly allow remote code execution (RCE) or unauthorized data exposure. However, it can disrupt production pipelines and microservices responsible for parsing incoming documents.
In cloud native or containerized environments, a crash of the document processing worker will cause container orchestration systems (like Kubernetes) to initiate a restart sequence. If the application pulls files from a queue and repeatedly attempts to parse the same malicious file upon restarting, the service can enter a continuous crash loop. This condition leads to queue starvation and broader system instability.
The CVSS v4.0 base score is calculated at 4.8 (Medium), with the vector CVSS:4.0/AV:L/AC:L/AT:N/PR:N/UI:P/VC:N/VI:N/VA:L/SC:N/SI:N/SA:N. This score reflects a local attack vector (as the file must be processed by the local host) with low overall availability impact under standard scoring models, though the operational impact in multi-tenant SaaS platforms can be significant.
The primary remediation strategy is upgrading the pypdf dependency to version 6.15.0 or later. This version enforces maximum limits on /ToUnicode token inputs and throws a LimitReachedError when encountering structured anomalies, preventing memory exhaustion.
pip install --upgrade pypdf>=6.15.0For environments where immediate upgrading is not possible, defensive engineers should apply sandboxing and process-level resource constraints. Applications handling PDF processing should isolate parser execution to worker nodes configured with memory limits using control groups (cgroups) or container-level specifications. Setting a hard memory limit ensures that a crash in a parser thread is isolated and does not affect the host node.
Furthermore, security audits should note a potential research gap in the fix. The current implementation strictly monitors token lengths in parse_bfrange. However, other CMap mapping structures, such as bfchar mappings (which translate single character codes), may remain unvalidated if they rely on a different code path inside _cmap.py. Security teams should monitor modifications to the library and apply validation layers globally on incoming stream sizes before handing them to the parser library.
CVSS:4.0/AV:L/AC:L/AT:N/PR:N/UI:P/VC:N/VI:N/VA:L/SC:N/SI:N/SA:N| Product | Affected Versions | Fixed Version |
|---|---|---|
pypdf py-pdf | < 6.15.0 | 6.15.0 |
| Attribute | Detail |
|---|---|
| CWE ID | CWE-400 |
| Attack Vector | Local (via crafted document parsing) |
| CVSS Score | 4.8 |
| Impact | Denial of Service (OOM Process Crash) |
| Exploit Status | Proof-of-Concept only |
| KEV Status | Not listed |
The product does not properly control the allocation and maintenance of a limited resource, enabling an actor to influence the amount of resources consumed and eventually leading to exhaustion.
An uncontrolled resource consumption vulnerability exists in the Docling document conversion library. Maliciously structured HTML, JATS, ODS, or BoxNote inputs containing table cells with excessively large 'rowspan' or 'colspan' attribute values trigger algorithmic complexity conditions. This allows unauthenticated remote attackers to initiate resource exhaustion states, crashing or hanging the target document processing pipeline while bypassing configured timeouts.
A Local File Inclusion (LFI) and Arbitrary File Disclosure vulnerability exists in Docling and Docling Slim versions >= 2.16.0 up to 2.131.0. When parsing serialized DoclingDocument structures using the JSON input format, the backend fails to restrict image URI schemes, allowing remote attackers to retrieve local files and verify path existence on the host system during embedded document export.
Docling, a tool for parsing and processing diverse document formats, is vulnerable to arbitrary file read, arbitrary file write, and potential remote code execution (RCE) in versions 2.94.0 through 2.131.0. The vulnerability occurs when applications configure Docling to use the Tectonic engine for rendering TikZ diagrams into images. Because the compilation did not restrict hazardous TeX primitives or sandbox the environment, an attacker can supply crafted documents containing malicious TikZ definitions to access or modify local files and execute arbitrary commands under the privileges of the processing application.
An SSRF guard bypass vulnerability in the Docling document conversion engine allows unauthenticated attackers to bypass internal IP access controls. The vulnerability exists due to a DNS rebinding Time-of-Check Time-of-Use (TOCTOU) condition, URL authority parsing inconsistencies, and unvalidated network requests triggered during headless browser page rendering.
A technical analysis of CVE-2026-105742 (GHSA-p3fw-7699-7926), a sensitive information disclosure vulnerability in the Docling document processing library. Vulnerable versions of Docling indiscriminately forward custom HTTP headers, such as authentication tokens, to arbitrary third-party origins and during cross-origin redirects while fetching remote image assets from untrusted HTML and EPUB documents.
CVE-2026-106121 is a Denial of Service (DoS) vulnerability in the RabbitMQ Java Client library (amqp-client) affecting versions prior to 5.37.0. The vulnerability resides in the legacy, custom JSON-RPC parsing class com.rabbitmq.tools.json.JSONReader. When parsing malformed or truncated payloads ending within a quoted string or single-line comment, the parser's scanner enters an infinite loop. This occurs because the loop lacks an exit condition for the end-of-input sentinel character returned by the iterator, leading to either CPU exhaustion or a JVM crash from an OutOfMemoryError.