Skip to content

DOCOPS-183 Add @neo4j-antora/pdf-generator: PDF export for Antora docsets - #136

Open
recrwplay wants to merge 22 commits into
devfrom
pdf-generator-package-clean
Open

recrwplay wants to merge 22 commits into
devfrom
pdf-generator-package-clean

Conversation

@recrwplay

@recrwplay recrwplay commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Introduces a real PDF export pipeline for any Antora docset (DOCOPS-183), replacing the earlier manual-checkout pilot (pdf-generation-pilot, now superseded and abandoned) with an installable npm package.

  • @neo4j-antora/pdf-generator - a new package bundling the print theme, convert.js (wraps asciidoctor-web-pdf), and the Antora Postprocessors that port roles-labels/table-footnotes for the PDF pipeline (they can't run as real Antora extensions there - see the package's own README).
    • Self-registers as an Antora extension with a sensible default config, so it needs no dedicated playbook or package.json script of its own - just pass it as a one-off flag onto whatever a docset already uses for a normal build: antora publish.yml --extension @neo4j-antora/pdf-generator, or npm run verify:publish -- --extension @neo4j-antora/pdf-generator reusing an existing script.
    • Solves the real version clash between Antora (@asciidoctor/core 2.x) and asciidoctor-pdf (needs 4.x) with its own isolated install, set up automatically by its own postinstall - no sibling npm project for a docset to know about.
    • Depends on the real, published @neo4j-antora/roles-labels@0.1.14 (Extract roles-labels' per-element transform into a shared, reusable core #134) so labels render identically between HTML and PDF - no hand-ported subset to drift out of sync.
    • The two Asciidoctor.js API incompatibilities that used to need patch-package (macros, remote-include) are now runtime shims instead - more robust, since they don't depend on install scripts running.
  • .github/workflows/reusable-docs-pdf-build.yml + docs-generate-pdf.yml - CI entry point for any docset, deliberately mirroring reusable-docs-build.yml's own shape: a package-script input (default verify:publish) trusts the docset to already have a valid script, the same convention and trust model reusable-docs-build.yml already uses - no dedicated PDF playbook/script needed anywhere. Installs dependencies, forces Antora to 3.2.0+ (needed for @antora/pdf-extension's assembler, regardless of what a docset's own package.json pins for its HTML build), runs the build with @neo4j-antora/pdf-generator injected via --extension (the same mechanism reusable-docs-build.yml already uses to inject @neo4j-antora/tabbed-nav), uploads the result as a build artifact.
    • Considered folding this directly into reusable-docs-build.yml itself, since checkout/install/antora-version/package-script are all already identical and the only real gap is finding+uploading the PDF artifact - decided against it, since that workflow is used across many other repos already and PDF generation is still young enough to want its own isolated blast radius (a PDF-pipeline failure should never risk the real site publish).
    • Generating a PDF this way always regenerates the full HTML site alongside it in the same Antora run (@antora/assembler hooks into normal site generation rather than replacing it) - an accepted, known build-time cost for now.

Known gaps (see pdf-generator/README.adoc)

  • Footnotes don't render at all in the PDF (root cause not yet found - upstream in asciidoctor-web-pdf/Vivliostyle, not this package).
  • mark-terms needs a product decision on how "first use" should work in a PDF (no per-page concept) before it can be wired in.

Before merging

  • @neo4j-antora/pdf-generator@0.1.0 needs publishing to npm (currently only prerelease tags 0.1.0-rc.1/rc.2 exist) - docs/package.json already points at the real 0.1.0.

Test plan

  • Verified end-to-end with a real npm install of the published prerelease package (cache cleared, not a symlink) into this exact branch state, running the exact command the workflow now runs (npm run verify:publish -- --extension @neo4j-antora/pdf-generator) - clean log (zero warn/error/fatal) and correct PDFs: role-based and inline labels, synonym/version-suffix resolution, and custom inline label text all render exactly as they do on the real HTML site.
  • CI checks pass

Introduces a real PDF export pipeline for Antora docsets (DOCOPS-183),
built as an installable npm package rather than a manual checkout
convention. @neo4j-antora/pdf-generator bundles the theme, convert.js,
postprocessors, and its own isolated Asciidoctor.js install (needed
because asciidoctor-pdf requires @asciidoctor/core 4.x while Antora
needs 2.x - handled automatically via the package's own postinstall,
not a sibling project a docset has to know about). It self-registers
as an Antora extension with a default config, so a docset's playbook
only needs:

  antora:
    extensions:
     - require: '@neo4j-antora/pdf-generator'

(also therefore usable directly as `antora --extension
@neo4j-antora/pdf-generator`). Depends on the real, published
@neo4j-antora/roles-labels@0.1.14 (#134) for label handling shared
identically with the HTML build - no hand-ported subset to drift out
of sync.

docs/pdf.yml is this repo's own test docset's playbook (content
sources/attributes differ per docset, so not shared, same as
preview.yml/publish.yml). docs/package.json adds pdf-generator as a
normal dependency and a build:pdf script.

.github/workflows/reusable-docs-pdf-build.yml and docs-generate-pdf.yml
give any docset a CI entry point: install dependencies (pulling in
pdf-generator and its own postinstall), force the Antora version to
3.2.0+ (needed for @antora/pdf-extension's assembler, regardless of
what a docset's own package.json pins for its HTML build), run the
PDF build, and upload the result as a build artifact.

Verified end-to-end: a real npm install of the published package
(cache cleared, not a symlink) into this exact branch state produces
a correct PDF - role-based and inline labels, synonym/version-suffix
resolution, custom inline label text all render exactly as they do on
the real HTML site.

Currently points at pdf-generator@0.1.0, not yet published (only
prerelease tags 0.1.0-rc.1/rc.2 exist) - needs publishing before this
lands.
@neo4j-antora/pdf-generator is a real, self-registering Antora
extension, so it doesn't need its own playbook file at all: passing
it as a one-off `antora <playbook> --extension @neo4j-antora/pdf-generator`
CLI flag works against any playbook a docset already has. Simpler for
docset authors (no new file to add or keep in sync) and avoids a
subtle duplication - generating a PDF always regenerates the full HTML
site alongside it in the same Antora run (assembler/pdf-extension hook
into the normal site-generation pipeline rather than replacing it), so
having a whole separate playbook for it was never buying independence
anyway.

- docs/pdf.yml removed; docs/package.json's build:pdf/verify:pdf
  scripts now run preview.yml with the --extension flag instead.
- reusable-docs-pdf-build.yml's pdf-playbook input renamed to
  antora-playbook, defaulting to publish.yml (what a real PDF export
  should reflect) instead of a dedicated pdf.yml default.
- Fixed a real bug this surfaced: the "find the PDF" step was
  hardcoded to build/site, which breaks for publish.yml (outputs to
  build/docs) - now searches build/ generally, excluding
  build/assembler/'s own intermediate copies specifically (which
  otherwise falsely match too).
- docs-generate-pdf.yml (this repo's own self-test) explicitly passes
  antora-playbook: preview.yml, since docs/ is a synthetic fixture and
  preview.yml (not publish.yml) is what actually registers tabbed-nav
  and carries the test content this pipeline has been verified
  against.

Deliberately keeping the PDF build as its own separate CI step rather
than folding --extension into the real HTML publish workflow, even
though that would avoid rebuilding the HTML twice - isolates the real
site publish from PDF-pipeline failures while it's still young
(footnotes are already a known-broken case).

Re-verified locally: `npm run build:pdf` against this exact branch
state (real npm install, prerelease package for now) produces correct
PDFs for both versions in this repo's multi-version test fixture.
Matches the convention already used when discussing this elsewhere
(docs-template's own build:pdf script) - a PDF export should reflect
published content, and this repo's own self-test docset doesn't need
a special case: the label/footnote test content lives in the same
source pages either playbook builds from, so publish.yml already
exercises it. Removes the now-redundant explicit antora-playbook
override in docs-generate-pdf.yml (matches reusable-docs-pdf-build.yml's
own default).

Re-verified locally against publish.yml: correct PDF output, and the
PDF-path lookup correctly finds it under build/docs/ (publish.yml's
own output.dir, different from preview.yml's build/site).
@recrwplay recrwplay changed the title Add @neo4j-antora/pdf-generator: PDF export for Antora docsets (DOCOPS-183) DOCOPS-183 Add @neo4j-antora/pdf-generator: PDF export for Antora docsets Sep 23, 2026
recrwplay and others added 19 commits September 24, 2026 11:56
reusable-docs-pdf-build.yml no longer takes an antora-playbook input
at all - it takes package-script (default verify:publish), same
convention and same trust model as reusable-docs-build.yml's own
input: the docset is trusted to already have a valid script there.
@neo4j-antora/pdf-generator is appended as a one-off `--extension`
flag via `npm run "$PACKAGE_SCRIPT" -- --extension ...`, the same
mechanism reusable-docs-build.yml already uses to inject
@neo4j-antora/tabbed-nav onto verify:publish.

This means a docset needs no dedicated PDF playbook *or* package.json
script at all - it just reuses whatever it already has. Removed
docs/package.json's now-redundant build:pdf/verify:pdf scripts
accordingly, and updated pdf-generator's own README to document the
--extension flag as the primary usage pattern (the declarative
playbook `require:` form still works and is documented as an
alternative).

Considered folding this directly into reusable-docs-build.yml itself
(since checkout/install/antora-version/package-script are all already
identical, and the only real gap is finding+uploading the PDF
artifact) but decided against it - reusable-docs-build.yml is used
across many other repos already, and keeping PDF generation as its
own separate, isolated workflow avoids adding untested complexity to
shared, widely-used infrastructure while this pipeline is still young.

Re-verified locally: `npm run verify:publish -- --extension
@neo4j-antora/pdf-generator` (the exact command the workflow now runs)
produces a clean log (zero warn/error/fatal) and a correct PDF.
…f-generator

package-script's own command (e.g. verify:publish) has no extensions
baked in itself - reusable-docs-build.yml supplies them as CLI flags
at call time, so calling that same script directly here was silently
running the HTML regenerated alongside the PDF (same Antora run, see
README) without roles-labels, xref-hash-validator, aliases-redirects,
etc. at all.

Adds antora-extensions (defaulting to the same list
reusable-docs-build.yml uses, plus @neo4j-antora/pdf-generator itself)
and antora-extensions-exclude, mirroring reusable-docs-build.yml's own
inputs and its "Remove excluded extensions" step verbatim - accepted
as a deliberate near-duplicate for now rather than trying to share the
list across both workflows, which would need either modifying the
widely-used reusable-docs-build.yml or a more complex nested-workflow-
call restructuring; not worth the risk/complexity for this.

docs-generate-pdf.yml (this repo's own self-test) now passes
antora-extensions-exclude for the three extensions this repo's own
docs/package.json doesn't actually install (antora-modify-sitemaps,
antora-page-list, antora-unlisted-pages) - discovered because running
the real default list locally throws Cannot find module for exactly
these three; same reason reusable-docs-build.yml's own exclude input
exists; no docset installs every extension in the full list.

Re-verified locally with the exact extension list and exclusions this
workflow now produces: clean log (only expected info-level messages
from roles-labels/table-footnotes/aliases-redirects/selector-labels
actually running) and a correct PDF.
@neo4j-antora/pdf-generator was listed in docs/package.json's regular
dependencies, which made every npm install against docs/ (including
the unrelated HTML PR-check build) require it to resolve. Remove it
from package.json and have reusable-docs-pdf-build.yml install it
directly via a new pdf-generator-version input, the same way it
already force-installs a specific Antora version.
Overriding asciidoctor-web-pdf's stylesheet attribute wholesale drops
its default theme entirely, including the span.footnote { float:
footnote } rule its footnote support depends on - without it, a
footnote's text just sits inline in the running text with no number
and no separation, easy to mistake for the macro being dropped. Add
the missing footnote CSS rules to print.css.
Now that span.footnote { float: footnote } is in place, the CSS
layout engine itself already places a table's footnotes on whatever
page that table lands on - the entire reason
table-footnotes-postprocessor.js existed. It only ever fired against
the old HTML5-backend #footnotes div structure, which
asciidoctor-web-pdf's own converter never produces, so it was always
a silent no-op here. Remove the postprocessor and its wiring in
convert.js, drop @neo4j-antora/table-footnotes from
reusable-docs-pdf-build.yml's default antora-extensions list (an
Antora pagesComposed hook this pipeline's isolated Asciidoctor
conversion never triggers), and update the README accordingly.
Third-party dependencies in package.json files are pinned to an exact version, not a range.
…ion defaults

Add a trigger-generate-pdf job to docs-trigger-builds.yml that dispatches
docs-generate-pdf.yml as its own separate run for the branch that was pushed, in parallel
with and independent of the HTML dispatch. Add this branch to the push trigger while the
PDF workflow is tested; that line must be removed before the PR is merged.

Remove the sitemaps, page list and unlisted pages extensions from the default extensions of
reusable-docs-pdf-build.yml, since they only apply to an HTML site, and drop the
antora-extensions-exclude from docs-generate-pdf.yml that existed to remove them again.
Redirects only apply to an HTML site, so the PDF build does not need the aliases-redirects
extension. The defaults are now roles-labels, selector-labels, xref-hash-validator and
pdf-generator.
The version selector labels only apply to an HTML site. The defaults are now roles-labels,
xref-hash-validator and pdf-generator.
Add a files whitelist to package.json. Without one, npm packed everything in the folder
that is not excluded by default, including the vendor/node_modules, vendor/package.json and
vendor/package-lock.json that the postinstall script creates, so publishing from a folder
where the package had been installed produced a 34 MB tarball with 13,000 files. The
whitelist gives the same 13 files (75 kB) from any folder.

Bump the version to 0.1.1, because 0.1.0 was published with those files, and move the
default pdf-generator-version of the PDF build to match.
The PDF reusable workflow did not keep the Antora log, so a failed PDF build left nothing to
inspect but the console. Add a Print Antora log step and an Upload Log artifact step that follow
reusable-docs-build.yml. The print step also runs when the PDF lookup fails, since the renderer
can fail without making the Antora run itself fail.
…ender, and bump to 0.1.2

A footnote reference inside a table that is itself nested in another table never finished
paginating in Vivliostyle: the renderer waited out its 180 s rendering timeout and the build
produced no PDF. Two cases were found with a 7 KB test document:

- an admonition (which Asciidoctor renders as a table) containing a table with a footnote.
  Lay the admonition out as plain blocks, so it is no longer a table in a table.
- a table inside a table cell (an a| cell holding a table) with a footnote.
  Do not float the footnote in that case; its text stays in the cell instead of moving to
  the foot of the page.

Footnotes in a plain paragraph, an admonition, a simple table and a table with header and
footer footnotes are unchanged. The docs-tools docs now build a PDF locally in a few seconds.

Bump the version to 0.1.2 and the default pdf-generator-version of the PDF build to match.
…tions, and bump to 0.1.3

The admonition title rules applied to every .title inside an admonition, not only its own title.
A table in an admonition has its caption as caption.title, so that caption was made
display: block, which laid it out after the header row (a thead is always laid out first)
instead of above the table, and shown in capitals in the admonition colour.

Scope the title rules to the admonition's own title, a direct child of its content cell. Table
captions in admonitions now look like those outside them.

Bump the version to 0.1.3 and the default pdf-generator-version of the PDF build to match.
@neo4j-docops-agent

Copy link
Copy Markdown
Collaborator

This PR includes documentation updates
View the updated docs at https://neo4j-docs-tools-136.surge.sh

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants