Skip to content

PDF detection requires a final %%EOF; rejects most real-world archived PDFs (17 of 20 in a msgvault batch) before upload #426

Description

@Joi

Summary

document/mistral.DetectFormat (via internal/formatdetect) rejects a PDF unless a %%EOF marker is present and is the final content of the file: PDF end marker is missing or not final. On a real mail archive that rule rejects most PDFs before anything is uploaded. In msgvault 0.20.0's first documents build over my archive, 17 of 20 standalone PDFs failed local preparation with invalid_local_source; all five I inspected fail this one check, and every one of them opens normally in Preview, pdf.js and Acrobat.

Environment

docbank v0.14.1-0.20260906150624-5969c0ba0811 as vendored by msgvault v0.20.0 (a36a394), darwin/arm64. Corpus: 231 PDF attachments from an Apple Mail archive spanning 1990s to 2026.

What the rejected files look like

Run through mistral.DetectFormat(f, size, "application/pdf") directly:

ok.pdf       size=63185  -> id="pdf"   (last %%EOF at 63178, followed by "\r\n")
failed.pdf   size=226095 -> err=PDF end marker is missing or not final  (no %%EOF anywhere; %PDF-1.3, 8 page objects, old Quartz producer)
6beaec4f.pdf size=231067 -> same  (no %%EOF anywhere; %PDF-1.4)
016f9e2a.pdf size=916418 -> same  (%PDF-1.6, only %%EOF is at offset 497 after the first-page xref; 915,916 bytes of objects follow, no final marker)
bdedccef.pdf size=233734 -> same  (%PDF-1.5, encrypted, only %%EOF at offset 499, 233,230 bytes follow)
942c7479.pdf size=9222   -> same  (%PDF-1.4, only %%EOF at offset 729, 8,488 bytes follow)

Two shapes: files with no end marker at all, and linearized or incrementally updated files whose only %%EOF belongs to the first cross-reference section. Both are common in the wild; PDF readers reconstruct the xref table and open them without complaint, and Mistral's OCR would almost certainly accept them.

What I expected

Fail-closed makes sense for "is this really a PDF", but the end marker is a weak signal for that: the %PDF- header plus a parseable page tree is the test readers actually use. Options, any of which would do:

  • Treat a missing or non-final %%EOF as a warning, and accept the file when the header is present and a lenient parse finds at least one page (pdfcpu's relaxed validation mode does this).
  • Search for the end marker within a bounded window of the tail rather than requiring it to be the final bytes, and accept a file whose only marker is early if a startxref or an object stream follows.
  • Expose a policy knob so an operator can opt into lenient PDF acceptance for archives like this one.

Whatever the choice, the msgvault side currently reports only invalid_local_source, so the operator cannot see which check failed; I am filing that separately against msgvault.

Reproduce

Take any PDF, strip the final %%EOF line (or use a linearized PDF from an older Acrobat), and call mistral.DetectFormat with the PDF media type. Happy to send the six digests or run a patched build against the corpus.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions