Repository navigation
Add preserve_whitespace option to Resolver - #13
Conversation
|
Thanks — the need makes sense, and skipping the squish is the useful half of this. The Line breaks land mid-sentence wherever bold or underline begins. ODS works because The simplest fix would be to drop the separator plumbing and let Whether that's enough depends on your use case: which format are you actually A few smaller things too (the constructor keyword vs the existing |
54c588b to
cb6a6f7
Compare
Resolver#text collapses all whitespace into single spaces, which is what search indexing wants but discards the structure the extraction commands emit. Consumers that process the text paragraph by paragraph have no way to keep it short of bypassing the Resolver. preserve_whitespace skips the squish and nothing else, so it yields the document structure for the formats whose handlers emit one: PDF, RTF and plain text. It is an accessor rather than a constructor argument, for consistency with max_plaintext_bytes. The zipped XML handlers join their text elements with a space and are unaffected. Making those emit real paragraph boundaries needs the separator keyed on the paragraph element rather than the run element, which is a separate change.
cb6a6f7 to
11607eb
Compare
|
You're right, and thanks for the detailed diagnosis. I checked the fixture: text.docx has a single w:p with five w:t runs, so the separator lands exactly where you said it would. The OOXML half of this was wrong. I've dropped the separator plumbing entirely. preserve_whitespace now only skips the squish, which is the useful half and is correct for the formats whose handlers already emit real structure. I also switched it to an accessor for consistency with max_plaintext_bytes, which covers your other two points, and rebased onto master. Added an end-to-end spec on the PDF fixture asserting the line structure survives in preserve mode and that the default output still has no line breaks at all. The PR description is updated too. On your question: PDF is the main format on our side, so this covers our case. Keying the separator on w:p for DOCX is worth doing properly, but I agree it belongs in its own change. |
The README listed only PDF, RTF and plain text as keeping their structure, but every format extracted by an external command does: catdoc, catppt, xls2csv and tesseract emit line breaks as well. It now also states that the text is returned as the handler emits it, form feeds between PDF pages, \r\n line endings and edge whitespace included. The preserve mode spec claimed to check composition but fed an already composed string, so it only exercised the byte limit. It now uses a decomposed character, which only fits the limit once composed.
|
thank you, that's merged now. |
Motivation
Resolver#textcollapses all whitespace into single spaces, which is what search indexing wants but discards the structure the extraction commands emit. Consumers that process the extracted text paragraph by paragraph have no way to keep it short of bypassing theResolverentirely.What this PR does
Adds a
preserve_whitespaceaccessor toPlaintext::Resolver. When set, the squish and strip are skipped and nothing else changes.That yields the document structure for the formats whose handlers already emit one: plain text and everything extracted by an external command (PDF, DOC, XLS, PPT, RTF and images). The text is returned as the handler emits it, so form feeds between PDF pages,
\r\nline endings and edge whitespace come through as well; the README says so. The zipped XML handlers join their text elements with a space and are unaffected, as discussed below.It is an accessor rather than a constructor argument, for consistency with the existing
max_plaintext_bytes.Backward compatibility
preserve_whitespacedefaults tofalse, so the output is unchanged unless it is set. No existing spec was modified; the spec changes are additions only:Not included
The earlier revision of this PR also keyed a
separator: "\n"throughZippedXmlHandler. As pointed out in review that was wrong:OfficeDocumentHandlercollectstelements, which are runs rather than paragraphs, so the separator landed mid sentence wherever formatting changed, andXlsxHandlerreads the deduplicated shared string table where separators bear no relation to rows or cells. That plumbing is gone.Making the zipped XML formats emit real paragraph boundaries needs the separator keyed on the paragraph element instead, which is a bigger change and better done on its own.