Summary
document/mistral.DetectFormat (via internal/formatdetect) rejects a PDF unless a %%EOF marker is present and is the final content of the file: PDF end marker is missing or not final. On a real mail archive that rule rejects most PDFs before anything is uploaded. In msgvault 0.20.0's first documents build over my archive, 17 of 20 standalone PDFs failed local preparation with invalid_local_source; all five I inspected fail this one check, and every one of them opens normally in Preview, pdf.js and Acrobat.
Environment
docbank v0.14.1-0.20260906150624-5969c0ba0811 as vendored by msgvault v0.20.0 (a36a394), darwin/arm64. Corpus: 231 PDF attachments from an Apple Mail archive spanning 1990s to 2026.
What the rejected files look like
Run through mistral.DetectFormat(f, size, "application/pdf") directly:
ok.pdf size=63185 -> id="pdf" (last %%EOF at 63178, followed by "\r\n")
failed.pdf size=226095 -> err=PDF end marker is missing or not final (no %%EOF anywhere; %PDF-1.3, 8 page objects, old Quartz producer)
6beaec4f.pdf size=231067 -> same (no %%EOF anywhere; %PDF-1.4)
016f9e2a.pdf size=916418 -> same (%PDF-1.6, only %%EOF is at offset 497 after the first-page xref; 915,916 bytes of objects follow, no final marker)
bdedccef.pdf size=233734 -> same (%PDF-1.5, encrypted, only %%EOF at offset 499, 233,230 bytes follow)
942c7479.pdf size=9222 -> same (%PDF-1.4, only %%EOF at offset 729, 8,488 bytes follow)
Two shapes: files with no end marker at all, and linearized or incrementally updated files whose only %%EOF belongs to the first cross-reference section. Both are common in the wild; PDF readers reconstruct the xref table and open them without complaint, and Mistral's OCR would almost certainly accept them.
What I expected
Fail-closed makes sense for "is this really a PDF", but the end marker is a weak signal for that: the %PDF- header plus a parseable page tree is the test readers actually use. Options, any of which would do:
- Treat a missing or non-final
%%EOF as a warning, and accept the file when the header is present and a lenient parse finds at least one page (pdfcpu's relaxed validation mode does this).
- Search for the end marker within a bounded window of the tail rather than requiring it to be the final bytes, and accept a file whose only marker is early if a
startxref or an object stream follows.
- Expose a policy knob so an operator can opt into lenient PDF acceptance for archives like this one.
Whatever the choice, the msgvault side currently reports only invalid_local_source, so the operator cannot see which check failed; I am filing that separately against msgvault.
Reproduce
Take any PDF, strip the final %%EOF line (or use a linearized PDF from an older Acrobat), and call mistral.DetectFormat with the PDF media type. Happy to send the six digests or run a patched build against the corpus.
Summary
document/mistral.DetectFormat(viainternal/formatdetect) rejects a PDF unless a%%EOFmarker is present and is the final content of the file:PDF end marker is missing or not final. On a real mail archive that rule rejects most PDFs before anything is uploaded. In msgvault 0.20.0's firstdocuments buildover my archive, 17 of 20 standalone PDFs failed local preparation withinvalid_local_source; all five I inspected fail this one check, and every one of them opens normally in Preview, pdf.js and Acrobat.Environment
docbank v0.14.1-0.20260906150624-5969c0ba0811 as vendored by msgvault v0.20.0 (a36a394), darwin/arm64. Corpus: 231 PDF attachments from an Apple Mail archive spanning 1990s to 2026.
What the rejected files look like
Run through
mistral.DetectFormat(f, size, "application/pdf")directly:Two shapes: files with no end marker at all, and linearized or incrementally updated files whose only
%%EOFbelongs to the first cross-reference section. Both are common in the wild; PDF readers reconstruct the xref table and open them without complaint, and Mistral's OCR would almost certainly accept them.What I expected
Fail-closed makes sense for "is this really a PDF", but the end marker is a weak signal for that: the
%PDF-header plus a parseable page tree is the test readers actually use. Options, any of which would do:%%EOFas a warning, and accept the file when the header is present and a lenient parse finds at least one page (pdfcpu's relaxed validation mode does this).startxrefor an object stream follows.Whatever the choice, the msgvault side currently reports only
invalid_local_source, so the operator cannot see which check failed; I am filing that separately against msgvault.Reproduce
Take any PDF, strip the final
%%EOFline (or use a linearized PDF from an older Acrobat), and callmistral.DetectFormatwith the PDF media type. Happy to send the six digests or run a patched build against the corpus.