ANT-2026-QT406EDT · libreoffice/core
other medium
Severity Claude high · Security research firm high · Maintainer medium
Discovered by Claude Mythos Preview
Anthropic's analysis, sealed at approval. Disclosure to the maintainer was performed by Ada Logics.
ANT-2026-QT406EDT: Heap overflow from unclamped /Length in hybrid-PDF stream extraction
In sdext/source/pdfimport/pdfparse/pdfentries.cxx, PDFObject::getDeflatedStream() allocates a heap buffer of nOuterStreamLen bytes (the lexically-measured distance between the 'stream' and 'endstream' tokens) but then sets the copy length to the raw dictionary /Length value via getDictLength() with no validation. The subsequent memmove at line 692 reads and writes that many bytes into the smaller allocation. This code path is reached from PDFDetector::detect() during default file-type detection whenever a PDF's trailer advertises /AdditionalStreams and a matching /DocChecksum — both of which the attacker authors and can trivially forge. An attacker who convinces a victim to open a crafted PDF therefore controls both the allocation size and the independent, larger copy length, yielding a heap OOB read+write that corrupts adjacent heap memory and is plausibly exploitable for code execution.
Target
Project: libreoffice/core
Location: sdext/source/pdfimport/pdfparse/pdfentries.cxx:692
Discovery: static analysis — not yet dynamically reproduced
Technical Details
The root cause is a missing bounds clamp: nOuterStreamLen (m_nEndOffset - m_nBeginOffset) is determined purely lexically by the boost::spirit grammar scanning for 'stream'/'endstream', while pBytes is set from getDictLength(), which returns the raw numeric /Length from the object dictionary unchecked. Because these two values are independent and attacker-controlled, memmove(rpStream.get(), pStream, pBytes) at line 692 can copy far more bytes than were allocated at line 653, producing a classic heap buffer overflow (both OOB read from and OOB write to adjacent heap memory).
Reproduction
- Author a PDF containing an object whose physical 'stream'...'endstream' body is small (e.g., 32 bytes) but whose dictionary declares a huge /Length (e.g., 1000000).
- Add a trailer with /AdditionalStreams referencing that object and a /DocChecksum equal to the MD5 of the file prefix (attacker computes this); omit new-style embedded files so getEmbeddedFile() fails and control falls through to getAdditionalStream().
- Deliver the PDF to the victim and have them open it in LibreOffice.
- PDFDetector::detect() → getAdditionalStream() → PDFObject::writeStream() → getDeflatedStream() allocates nOuterStreamLen bytes then memmoves /Length bytes, overflowing the heap.
[No reproducer or sanitizer output attached — request from security-cvd@anthropic.com if needed.]
Suggested Fix
Clamp the number of bytes copied and subsequently processed to the actual allocated buffer size (nOuterStreamLen) rather than trusting the file-declared /Length; reject or truncate streams whose /Length exceeds the lexical stream span.
Acknowledgement
This vulnerability was discovered by Claude, Anthropic's AI assistant, and triaged by the Anthropic security team in collaboration with Anthropic Research. Please direct questions to security-cvd@anthropic.com and reference ANT-2026-QT406EDT.
Reference: ANT-2026-QT406EDT
Anthropic CVD Policy: https://www.anthropic.com/coordinated-vulnerability-disclosure
Triage and disclosure were performed by Ada Logics.
- Verdict
- true positive
- Severity
- high
The change that resolved this finding.
diff --git a/sdext/source/pdfimport/pdfparse/pdfentries.cxx b/sdext/source/pdfimport/pdfparse/pdfentries.cxx
index d000a18dcea33..3ec950e22b755 100644
--- a/sdext/source/pdfimport/pdfparse/pdfentries.cxx
+++ b/sdext/source/pdfimport/pdfparse/pdfentries.cxx
@@ -688,6 +688,12 @@ bool PDFObject::getDeflatedStream( std::unique_ptr<char[]>& rpStream, unsigned i
pStream++;
// get the compressed length
*pBytes = m_pStream->getDictLength( pObjectContainer );
+ unsigned int nAvailable = nOuterStreamLen - static_cast<unsigned int>(pStream - rpStream.get());
+ if (*pBytes > nAvailable)
+ {
+ SAL_WARN("sdext.pdfimport.pdfparse", "stream /Length " << *pBytes << " exceeds " << nAvailable << " available bytes");
+ *pBytes = nAvailable;
+ }
if( pStream != rpStream.get() )
memmove( rpStream.get(), pStream, *pBytes );
if( rContext.m_bDecrypt )https://github.com/LibreOffice/core/commit/dcf16d610a32861349d4e4b84e7bef4b9420929c
Recorded dates, in order.
- 2026-04-02 Discovered or logged
- 2026-07-04 Sent to maintainer
- 2026-07-24 Patch released
- 2026-08-12 Maintainer acknowledged
- 2026-09-28 Publicly revealed
SHA-3-512 hash:
a546a862caaa6a6280b9ca7f696d78567d009cca5ef139a0b61306b5bdbdb7c494c58aac26ce4f0314b03512abf5ee240d50af1f8905096abda281f03e6e163e
Committed 2026-07-22 07:32 UTC
Revealed 2026-09-28 20:49 UTC
Verify (download preimage.json)
Show preimage JSON
{
"ant_id": "ANT-2026-QT406EDT",
"bug_class": "Memory corruption (heap buffer overflow)",
"claude_severity": "high",
"commit_sha": null,
"created_at": "2026-04-16T01:54:44+00:00",
"description": "In sdext/source/pdfimport/pdfparse/pdfentries.cxx, PDFObject::getDeflatedStream() allocates a heap buffer of nOuterStreamLen bytes (the lexically-measured distance between the 'stream' and 'endstream' tokens) but then sets the copy length to the raw dictionary /Length value via getDictLength() with no validation. The subsequent memmove at line 692 reads and writes that many bytes into the smaller allocation. This code path is reached from PDFDetector::detect() during default file-type detection whenever a PDF's trailer advertises /AdditionalStreams and a matching /DocChecksum — both of which the attacker authors and can trivially forge. An attacker who convinces a victim to open a crafted PDF therefore controls both the allocation size and the independent, larger copy length, yielding a heap OOB read+write that corrupts adjacent heap memory and is plausibly exploitable for code execution.",
"discovered_at": "2026-04-02T00:00:00+00:00",
"location": "sdext/source/pdfimport/pdfparse/pdfentries.cxx:692",
"poc_sha256": null,
"preimage_version": 1,
"project": "LibreOffice/core",
"reproduction": [
"1. Author a PDF containing an object whose physical 'stream'...'endstream' body is small (e.g., 32 bytes) but whose dictionary declares a huge /Length (e.g., 1000000).",
"2. Add a trailer with /AdditionalStreams referencing that object and a /DocChecksum equal to the MD5 of the file prefix (attacker computes this); omit new-style embedded files so getEmbeddedFile() fails and control falls through to getAdditionalStream().",
"3. Deliver the PDF to the victim and have them open it in LibreOffice.",
"4. PDFDetector::detect() → getAdditionalStream() → PDFObject::writeStream() → getDeflatedStream() allocates nOuterStreamLen bytes then memmoves /Length bytes, overflowing the heap."
],
"technical_details": "The root cause is a missing bounds clamp: nOuterStreamLen (m_nEndOffset - m_nBeginOffset) is determined purely lexically by the boost::spirit grammar scanning for 'stream'/'endstream', while *pBytes is set from getDictLength(), which returns the raw numeric /Length from the object dictionary unchecked. Because these two values are independent and attacker-controlled, memmove(rpStream.get(), pStream, *pBytes) at line 692 can copy far more bytes than were allocated at line 653, producing a classic heap buffer overflow (both OOB read from and OOB write to adjacent heap memory).",
"title": "Heap overflow from unclamped /Length in hybrid-PDF stream extraction",
"vendor_severity": "high"
}