A linked data publication pipeline that transforms a curated Zotero library into a knowledge graph and static website. The collection documents artists' books held at the Joseph C. Sloane Art Library (UNC Chapel Hill), along with citation data tracking how those books appear in reference literature.
The pipeline runs as two parallel tracks that meet at the graph stage: the artists' books (the 1,341-book collection the site is built from) and the reference works that cite them (the 157-item Reference resources collection, plus the "Cited:" notes connecting the two). Both are surfaced on the website and cross-linked — a book lists the reference works citing it, a reference work lists the books it cites. Only reference works that actually cite an in-collection book get a page (43 of the 157 today).
See docs/PIPELINE.md for a useful diagram, and
docs/README.md for an index of the supplementary
documentation.
The stages, and where each is documented in depth:
- Zotero → CSV/notes.
dedup.pycollapses the three Zotero libraries into the canonical book (~7.9k) and reference-work lists;notes_export.shemits the "Cited:" note paragraphs; andfreeze_citations.pyreconciles those paragraphs to reference works and freezes the citation edges intosources/citations.ttl. Both fuzzy steps are gated by a hand-owned decisions overlay — a person'ssame/no/unsurecalls insources/artists-books-dedup-decisions.csv(per lib-1 key) andyes/no/unsurecalls insources/citations-decisions.csv(per reference-work × book pair). These are read back but never rewritten bymake, so a re-run never clobbers them; each step also emits a generated*-review.csvsurfacing the uncertain matches to curate. →sources/zotero/README.md(database, notes, citation data model),sources/README.md(canonical dedup + frozen citations). - Library catalogs → MARCXML.
marc_harvest.py(YAZ) harvests full MARC records over Z39.50/SRU — books from UNC's catalog, reference works from a chain of nine catalogs by ISBN then title+author — each stamped with a synthetic999 $a <itemKey>to join back to its Zotero item. →sources/marc/README.md(harvest),docs/MARC-RECORDS.md(field-level analysis). - CSV + MARCXML → RDF graph. Two
SPARQL-Anything CONSTRUCT queries
(
queries/artists-books.rq,queries/reference-works.rq) emit BIBFRAME plus a customab:namespace intograph/*.ttl. The citation edges are not constructed here — they are the frozensources/citations.ttl(issue #82): 5,022 citations link 3,896 books to 54 reference works. Two parallel local SKOS vocabularies are mined from the book MARC and curated by hand, each gated by its owndecisions.csvoverlay (include? / concept / category, per heading cluster; hand-edited, never rewritten bymake): construction techniques, materials, binding/format types and printing methods (how a book is made) insources/construction/decisions.csv→sources/construction-methods.ttl, and topical/geographic subjects (what it is about) insources/subjects/decisions.csv→sources/subject-terms.ttl. Theirbuild-scheme.rqqueries emit the schemes, andartists-books.rqattaches anab:constructedUsing/ab:hasSubjectlink from each book to the concepts its headings map to. →CLAUDE.md(query architecture, URI minting),sources/construction/README.mdandsources/subjects/README.md(the mine → curate → build vocabulary pipeline),docs/QUERY-PERFORMANCE.md. - RDF graph → website. Apache
Fuseki loads the two
constructed graphs, the frozen citations, and the
construction-methods.ttlandsubject-terms.ttlSKOS schemes, and serves a local SPARQL endpoint; Snowman runs the SELECT queries inweb/queries/against it and renders the Go templates inweb/templates/intoweb/site/— including, on each book page, the construction-methods and subject sections resolved from the schemes. →CLAUDE.md(views, templates, Snowman gotchas).
make all # build graph/*.ttl, then web/site/index.html
make serve # build if needed, then serve at http://127.0.0.1:8080All tools are fetched (and, for YAZ, built from source) on first use —
no third-party binaries need a system-wide install. The only system
prerequisites are a JVM, sqlite3, make, and a C toolchain (to
compile the vendored
YAZ); GitHub
Codespaces gets these via .devcontainer/devcontainer.json. See
CLAUDE.md for the full target list and build details.
Regenerating the inputs under sources/ additionally needs GNU Make
≥ 4.3, for grouped targets. macOS ships GNU Make 3.81, so there use
gmake (brew install make) — make -C sources … stops with an
explanatory error rather than misbuilding. The make all / make serve build above has no such requirement.
| Tool | Version | Purpose |
|---|---|---|
| Apache Jena | 6.1.0 | RDF validation, RDFS reasoning, graph diffing (riot, arq, shacl) |
| Apache Fuseki | 6.1.0 | In-process SPARQL endpoint server |
| SPARQL-Anything | 1.1.0 | CSV/MARCXML-to-RDF transformation via SPARQL CONSTRUCT |
| Snowman | 0.8.0 | SPARQL-driven static site generator |
| YAZ | 5.37.3 | yaz-client/yaz-marcdump — Z39.50/SRU MARC harvest; built from source |
The graphs emit BIBFRAME with a custom ab: namespace layered on top
(ab:ArtistsBook, ab:ReferenceWork, ab:Citation,
ab:cites/ab:citedBy, ab:constructedUsing, ab:hasSubject,
creator-role properties). The concepts that ab:constructedUsing and
ab:hasSubject point at live in companion SKOS schemes,
sources/construction-methods.ttl and sources/subject-terms.ttl —
mined from the book MARC and curated by hand (see
sources/construction/README.md and
sources/subjects/README.md).
docs/vocab.ttl
defines the vocabulary and docs/description.ttl is a hand-written
worked example (Ed Ruscha's Twentysix Gasoline Stations); make validate runs Jena's validator over both. Some terms are defined but
not yet emitted, and the citation terms still carry a legacy ex:
prefix pending normalization — see docs/README.md
and CLAUDE.md for current status.
The Zotero library encodes the citation data three ways — tags,
"Cited:" notes (~7,649), and dc:relation links (~2,366); the
notes are the form the pipeline extracts. See
sources/zotero/README.md → Citation
data for the full model.