Skip to content

Stop getSemanticHTML from turning every space into   - #4827

Closed
afonsojanu wants to merge 1 commit into
slab:mainfrom
afonsojanu:fix/semantic-html-single-space-preserved
Closed

afonsojanu wants to merge 1 commit into
slab:mainfrom
afonsojanu:fix/semantic-html-single-space-preserved

Conversation

@afonsojanu

Copy link
Copy Markdown

Fixes #4509.

convertHTML() in editor.ts replaces every space in a text blot with  , which is why getSemanticHTML() comes back with a non-breaking space between every single word instead of the plain spaces that were actually typed. That changes the meaning of the text, which is exactly what the issue is complaining about.

The reason the nbsp substitution exists at all is that HTML collapses a run of consecutive whitespace down to one visible character, so if you have two or three spaces in a row you do need at least one nbsp to keep them all visible. A single space between two words has never had that problem, so there's no reason to touch it.

Changed the replace to only kick in on runs of two or more consecutive spaces, keeping one of them as a plain space and turning the rest into nbsp. A lone space is left untouched.

I couldn't get the project's own browser-based test suite (Playwright/Chromium via vitest) to actually launch in the environment I was working in - Chromium installs fine but the browser process gets killed immediately on launch, unrelated to anything in this change. I checked the transform logic directly with a small standalone script instead and it does what's expected:

"a text with multiple spaces" -> "a text with multiple spaces"
"a  double space"             -> "a  double space"
"a   triple"                  -> "a   triple"
" leading"                    -> " leading"
"trailing "                   -> "trailing "

I also added two cases to the existing getSemanticHTML() describe block in quill.spec.ts covering the single-space and multi-space scenarios, written in the same style as the tests already there, but wasn't able to confirm they pass against the real suite for the reason above. Happy to double check that part if someone can run it against the actual browser suite.

convertHTML replaced every single space in a text blot with  ,
which changes the actual meaning of the text: "a text with multiple
spaces" came back with a non-breaking space between every word
instead of the plain spaces that were actually typed.

A run of two or more consecutive spaces genuinely needs at least one
non-breaking space, since HTML would otherwise collapse the whole run
down to a single visible space when rendered. A single space between
two words doesn't have that problem and should stay as a plain space.

Fixes #4509.
@afonsojanu

Copy link
Copy Markdown
Author

Closing this one too, same reason: starting fresh and clearing out everything I had open.

@afonsojanu afonsojanu closed this Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[2.0.3] getSemanticHTML is broken

1 participant