Summary
A hybrid-reference PDF cannot expose its page list when an unknown stream filter affects only an incremental-update object stream containing optional outline objects. Hybrid cross-reference parsing falls back to object scanning, which eagerly expands that unrelated object stream and aborts before the original page tree can be used.
Versions tested
Test files
Reproduction
import playa
with playa.open("UnknownFilter-OutlineObjStm.pdf") as document:
print(len(document.pages))
print(document.pages[0].extract_text())
print(len(list(document.pages[0].images)))
Actual behavior
Accessing document.pages first reports a hybrid cross-reference parsing failure, enters fallback object scanning, and then aborts while eagerly expanding the outline object stream:
NotImplementedError: Unsupported filter: /'XXXDecode'
Page enumeration, text extraction, and image extraction never run.
Expected behavior
The unknown filter affects only the object stream containing three optional outline objects. The existing page tree and its ordinary contents should remain available for page enumeration, text extraction, and image extraction; the outline itself may be unavailable with a warning or scoped error.
Relevant implementation
Document._read_xrefs() enters XRefFallback after hybrid cross-reference parsing raises PDFSyntaxError. XRefFallback._load() then expands every discovered /ObjStm immediately via stream.buffer, so an unsupported filter in an optional object stream aborts reconstruction globally.
Disclosure
This issue was prepared and submitted by an AI coding agent on behalf of @lin-string. The reproduction and results were verified against the current default branch before submission.
Summary
A hybrid-reference PDF cannot expose its page list when an unknown stream filter affects only an incremental-update object stream containing optional outline objects. Hybrid cross-reference parsing falls back to object scanning, which eagerly expands that unrelated object stream and aborts before the original page tree can be used.
Versions tested
playa-pdf1.1.0 (v1.1.0, commit85a9c1e22e327bf873e4282d66ed1dba89fa7575)mainat commit9496dfea5150343c9dc05544d9c004bbcccf17e6, checked on 2026-09-06Test files
Reproduction
Actual behavior
Accessing
document.pagesfirst reports a hybrid cross-reference parsing failure, enters fallback object scanning, and then aborts while eagerly expanding the outline object stream:Page enumeration, text extraction, and image extraction never run.
Expected behavior
The unknown filter affects only the object stream containing three optional outline objects. The existing page tree and its ordinary contents should remain available for page enumeration, text extraction, and image extraction; the outline itself may be unavailable with a warning or scoped error.
Relevant implementation
Document._read_xrefs()entersXRefFallbackafter hybrid cross-reference parsing raisesPDFSyntaxError.XRefFallback._load()then expands every discovered/ObjStmimmediately viastream.buffer, so an unsupported filter in an optional object stream aborts reconstruction globally.Disclosure
This issue was prepared and submitted by an AI coding agent on behalf of @lin-string. The reproduction and results were verified against the current default branch before submission.