Skip to content

Add preserve_whitespace option to Resolver - #13

Merged
jkraemer merged 2 commits into
planio-gmbh:masterfrom
bukhr:feat/preserve-whitespace-option
Sep 21, 2026
Merged

jkraemer merged 2 commits into
planio-gmbh:masterfrom
bukhr:feat/preserve-whitespace-option

Conversation

@diegocostares

@diegocostares diegocostares commented Jul 16, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Resolver#text collapses all whitespace into single spaces, which is what search indexing wants but discards the structure the extraction commands emit. Consumers that process the extracted text paragraph by paragraph have no way to keep it short of bypassing the Resolver entirely.

What this PR does

Adds a preserve_whitespace accessor to Plaintext::Resolver. When set, the squish and strip are skipped and nothing else changes.

That yields the document structure for the formats whose handlers already emit one: plain text and everything extracted by an external command (PDF, DOC, XLS, PPT, RTF and images). The text is returned as the handler emits it, so form feeds between PDF pages, \r\n line endings and edge whitespace come through as well; the README says so. The zipped XML handlers join their text elements with a space and are unaffected, as discussed below.

It is an accessor rather than a constructor argument, for consistency with the existing max_plaintext_bytes.

Backward compatibility

preserve_whitespace defaults to false, so the output is unchanged unless it is set. No existing spec was modified; the spec changes are additions only:

  • two resolver examples, one for the preserved whitespace and one asserting that composition and the byte limit still apply in preserve mode
  • an end-to-end example on the PDF fixture, asserting the line structure survives in preserve mode and that the default output contains no line breaks at all

Not included

The earlier revision of this PR also keyed a separator: "\n" through ZippedXmlHandler. As pointed out in review that was wrong: OfficeDocumentHandler collects t elements, which are runs rather than paragraphs, so the separator landed mid sentence wherever formatting changed, and XlsxHandler reads the deduplicated shared string table where separators bear no relation to rows or cells. That plumbing is gone.

Making the zipped XML formats emit real paragraph boundaries needs the separator keyed on the paragraph element instead, which is a bigger change and better done on its own.

@jkraemer

jkraemer commented Aug 3, 2026

Copy link
Copy Markdown
Member

Thanks — the need makes sense, and skipping the squish is the useful half of this.

The separator: "\n" half doesn't do what it says for the OOXML formats though.
DocxHandler's SAX element is t, which is a run, not a paragraph — runs split at
every formatting change. On the existing fixture:

default:  "lorem ipsum fulltext find me!"
preserve: "lorem ipsum \nfulltext\n \nfind\n me!\n"

Line breaks land mid-sentence wherever bold or underline begins. XlsxHandler is worse:
it reads xl/sharedStrings.xml, the deduplicated string table, so separators there bear
no relation to rows or cells, and a string used in ten cells appears once.

ODS works because OpendocumentHandler sets @element = 'p', which really is a
paragraph — and it's the only format the new specs cover.

The simplest fix would be to drop the separator plumbing and let preserve_whitespace
mean only "don't squish". That's correct for PDF, RTF and plain text, where the handler
already emits real structure.

Whether that's enough depends on your use case: which format are you actually
extracting?
For ODT or PDF it solves the problem. For DOCX it doesn't — that needs the
separator keyed on w:p rather than w:t, which is a bigger change worth doing properly.

A few smaller things too (the constructor keyword vs the existing max_plaintext_bytes
accessor, DEFAULT_SEPARATOR being applied in three places) — worth sorting once the
above is settled.

@diegocostares
diegocostares force-pushed the feat/preserve-whitespace-option branch from 54c588b to cb6a6f7 Compare August 5, 2026 22:30
Resolver#text collapses all whitespace into single spaces, which is what
search indexing wants but discards the structure the extraction commands
emit. Consumers that process the text paragraph by paragraph have no way
to keep it short of bypassing the Resolver.

preserve_whitespace skips the squish and nothing else, so it yields the
document structure for the formats whose handlers emit one: PDF, RTF and
plain text. It is an accessor rather than a constructor argument, for
consistency with max_plaintext_bytes.

The zipped XML handlers join their text elements with a space and are
unaffected. Making those emit real paragraph boundaries needs the
separator keyed on the paragraph element rather than the run element,
which is a separate change.
@diegocostares
diegocostares force-pushed the feat/preserve-whitespace-option branch from cb6a6f7 to 11607eb Compare August 6, 2026 04:42
@diegocostares

Copy link
Copy Markdown
Contributor Author

You're right, and thanks for the detailed diagnosis. I checked the fixture: text.docx has a single w:p with five w:t runs, so the separator lands exactly where you said it would. The OOXML half of this was wrong.

I've dropped the separator plumbing entirely. preserve_whitespace now only skips the squish, which is the useful half and is correct for the formats whose handlers already emit real structure. I also switched it to an accessor for consistency with max_plaintext_bytes, which covers your other two points, and rebased onto master.

Added an end-to-end spec on the PDF fixture asserting the line structure survives in preserve mode and that the default output still has no line breaks at all. The PR description is updated too.

On your question: PDF is the main format on our side, so this covers our case. Keying the separator on w:p for DOCX is worth doing properly, but I agree it belongs in its own change.

@diegocostares

Copy link
Copy Markdown
Contributor Author

@jkraemer 👀

The README listed only PDF, RTF and plain text as keeping their
structure, but every format extracted by an external command does:
catdoc, catppt, xls2csv and tesseract emit line breaks as well. It now
also states that the text is returned as the handler emits it, form
feeds between PDF pages, \r\n line endings and edge whitespace included.

The preserve mode spec claimed to check composition but fed an already
composed string, so it only exercised the byte limit. It now uses a
decomposed character, which only fits the limit once composed.
@jkraemer
jkraemer merged commit 34561f9 into planio-gmbh:master Sep 21, 2026
6 checks passed
@jkraemer

Copy link
Copy Markdown
Member

thank you, that's merged now.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants