This defines the scaffolding for turning a biographical source (a paper, an oral history, an archive) into a small, queryable knowledge graph + timeline. It is deliberately narrow: four JSON files per subject, each validated against a JSON Schema in this folder. The goal is a model tight enough that two different people extracting from the same paper would produce essentially the same data.
A subject is a person. Each source document about them gets its own
folder beneath, named by that source's citation key (sources[].id):
subjects/materials_science/ the collection this person was found for
suntola/
subject.json the person: canonical name, slug, domains[]
puurunen_2014/ one document's extraction -- the four files below
entities.json
events.json
relations.json
sources.json
aris_2019/ a second account of the same life
...
Each document folder is extracted independently and kept whole. Two papers about the same person are two accounts, not one merged truth — reconciling them (same human, different entity ids, the same episode described differently) happens when a subject is read, so a wrong reconciliation is a rendering bug you re-run rather than extraction you have destroyed. Where two accounts disagree, both survive with their own citations, and which document asserted what stays answerable.
On read, entities sharing an id are treated as the same thing and their aliases unioned; events and relations are never merged, and an id colliding across documents is prefixed with its document key.
subject.json carries exactly one canonical name for the person — the
fullest form seen — with every other form recorded in that person's
aliases, so the same scientist never appears under several spellings.
The rule above pulls name variants together. The opposite case is real and pulls the other way: two different people can share a name. Finland has two Pekka Soininens, both of whom worked at Microchemistry, and both matter to the history of atomic layer deposition — a fact pointed out by the author of the Suntola biography, not by anything in the documents.
Nothing detects this automatically, and merge-on-read makes it silent: two
documents that each produce pekka_soininen for two different men become one
node whose events belong to neither. So when a document distinguishes two people
of the same name, give them separate ids qualified by whatever distinguishes
them — pekka_soininen_beneq, pekka_soininen_microchemistry — and say which
is which in each summary. Split only when the document itself distinguishes
them; splitting on a guess is its own error.
The build prints a note when two documents describe one person id with nothing
in common, which is a prompt to look rather than a detector — it cannot catch
the Soininen case, where both descriptions would read "worked at Microchemistry".
It also warns when one person is the subject of two birth or death events,
which is the strong signal: a life has one of each.
entities.json— the nouns: people, places, organizations, and named "artifacts" (inventions, patents, products, publications, companies-as- things-founded). Everything an event or relation can point to.events.json— the things that happened. See "What counts as an event" below.relations.json— simple typed edges between two entities (visited,worked_at,lived_in, ...). Some relations are just a durable summary of an event (e.g. anemployment_startevent implies aworked_atrelation); others stand alone when the source states a fact without a datable event behind it (e.g. "they were lifelong friends").sources.json— the source document(s) this subject's data was extracted from. Every event and relation cites at least one source here.
A subject is self-contained: nothing in one subject's folder refers to an id
in another's. Cross-subject connections (the whole reason for
having two biographies) get their own bridge file once both subjects exist:
subjects/_bridges/<a>-<b>.json, using the same relation shape but with
entity ids qualified as subject:id (e.g. suntola:tuomo_suntola). Bridge
ids stay subject-scoped, not document-scoped: they connect two people, and
which document happened to mention the link is already recorded in that
relation's own sources[].
An event is a dateable occurrence, not a fact or a description. The test: could you point to roughly when it happened, even fuzzily? If the source only tells you a static fact with no timeframe ("Suntola was interested in eastern philosophy"), it is not an event — it's either dropped, or captured as a note on the entity.
Every event must have:
- A date, however imprecise (a specific day, a month, a year, a decade, or an explicitly approximate date) — never "undated."
- A type, drawn from the controlled vocabulary in
event.schema.json(event_type). If nothing fits, use"other"and explain in the description rather than inventing a new type ad hoc. - At least one participant — the entity or entities the event is
about. A "company founded" event participates the founder(s) and the
organization founded (with
rolesdistinguishing them). - At least one source citation —
source_idand apageare both required: provenance means a reader can turn to exactly where in the text this came from. A nonempty supportingquoteis required for every citation in a new extraction, including relation citations. The stored schemas accept missing quotes for legacy data only;extraction_schemas()adds the new-output requirement to the prompt, andcheck_extraction_quotes()rejects missing or blank quotes before new output is written. No event is entered on inference or general knowledge; if the paper doesn't say it, it doesn't go in.
Events are the join between the graph and the timeline: the timeline is
just events.json sorted by date; the graph is entities.json +
relations.json, and relations may (optionally) point back to the event
that established them via event_id, so clicking a graph edge can jump to
its moment in the timeline and vice versa.
Biographical sources rarely give clean ISO dates ("in the early 1970s", "by the end of 1973", "1963"). Rather than fake precision, every event date is an object with:
display— the human-readable string as it should be shown ("1963", "August–September 1974", "c. 1958", "early 1970s").precision— one ofday | month | season | year | decade | century | circa | range.seasonis for sources that date something to a season rather than a month ("summer 1964"), common in recollection and correspondence;monthwould invent a precision the source never gave.centuryexists because a collection reaching back before modern record-keeping needs it: a document on Jabir ibn Hayyan can date his father's Abbasid patronage to the 8th century and no further, and the alternative was discarding the whole extraction rather than saying so.sort_start/sort_end— ISOYYYY-MM-DDbounds used only for sorting and for drawing a point-vs-span on the timeline; they are the normalized uncertainty bounds, not an implied exact date. Equal bounds are reserved for an explicitly known single day.
The authoritative normalization policy is x-normalization-policy in
date.schema.json, included in every extraction prompt. Named calendar units
use their full bounds. Early/mid/late qualifiers stay in display without
invented cutoffs: "early June 1974" uses June 1–30. Approximate dates with no
explicit uncertainty interval expand by one named unit on each side as a
sorting convention, not an asserted historical limit. The policy also defines
seasons, explicit intervals and relative-date anchors. Existing stored dates
are not rewritten when this policy changes.
This keeps ordering well-defined (sort by sort_start, break ties by
sort_end) without ever putting a false-precision date in front of a
reader. A range-precision event (e.g. "worked at Lohja, 1974–1987")
renders as a bar on the timeline instead of a point.
entity_type is one of person | place | organization | artifact.
artifact covers inventions, patents, products, publications, and other
named things-that-aren't-people-places-or-orgs (e.g. "Atomic Layer
Epitaxy", "the 1974 ALE patent", "Humicap sensor"). Keeping these as
entities (not just event descriptions) lets the graph show, e.g., one
invention connected to every person who worked on it and every event in
its history.
Coordinates for the map view: a place entity that should appear on
the frontend's map view carries attributes.lat / attributes.lng
(decimal degrees) and, for provenance, attributes.wikidata_qid. Get
these from Wikidata's P625 (coordinate location) the same way as a
portrait (see "Portraits" below) — verify the candidate is the right
place (country/description match) before trusting its coordinates. A
place entity without lat/lng simply doesn't appear on the map; it
still works everywhere else (graph, timeline).
A relation is { source, type, target } plus optional time bounds
(start/end, same fuzzy-date shape as events, precision-only, no
display/sort split needed since relations aren't independently plotted)
and source citations. type is drawn from the controlled vocabulary in
relation.schema.json. Relations are directed; the vocabulary defines the
reading direction (e.g. worked_at: person → organization) — and
scripts/build_site.py's RELATION_DIRECTIONS table enforces it at
build time, checking that source/target actually have the expected
entity_type for that relation type, not just that the ids resolve.
Keep that table in sync with this vocabulary the same way as
event_type (see "Extending the vocabularies" below).
RELATION_MEANINGS additionally defines semantic direction, such as
supervisee → supervisor for supervised_by and mentor → mentee for mentored.
Both tables are embedded into the extraction schema's type description;
prompt construction fails if either table differs from the relation enum.
Only events have an other fallback. Omit an unrepresentable relation while
retaining any independently supported event. An explicit organization rename
uses old-name → new-name nodes as the exception to ordinary alias merging.
A slot may legitimately allow more than one entity_type. visited and
lived_in accept place or organization, because an institution is
honestly both a body and somewhere you can be: a college, a monastery, a
hospital. employed_by, licensed_to, acquired_by and sold_to accept a
person at either end — an apprentice is employed by a master, and individuals
really did hold licences and buy businesses, especially before incorporation
was common. founded accepts an artifact target, since founding a journal
is a real act.
None of that is a loophole for sloppy typing. Each was widened from a real
observed case, not from guesswork: Gresham College typed as an
organization and "Hooke lived there" are each correct, and a validator that
rejected the pair discarded an otherwise sound extraction over one edge. When
adding a relation type, ask whether its ends are genuinely single-typed before
constraining them.
worked_at is deliberately not widened to accept a place. Sources
routinely write "he went to Uppsala" meaning the university, but "worked at
Uppsala the city" and "worked at Uppsala University" are different claims, and
collapsing them would erode the person → organization guarantee that makes the
graph queryable. Such relations are set aside instead (below).
A relation that fails the direction check, or names an entity that doesn't
exist, is dropped from the build and written to
subjects/<slug>/<doc>/relations.rejected.json with its reason and citations
— rather than failing the whole subject. A relation is a leaf: nothing else
references it, so removing one leaves nothing else inconsistent. Dangling
references in events stay fatal, since events carry the timeline the rest of
the graph hangs off.
The parked files are worth reading periodically. A set-aside relation is often a variant reading rather than a mistake, so they are the evidence base for deciding which entries here are still too narrow.
All ids are lowercase snake_case, unique within their file's type and
stable once assigned (other files reference them). Prefer human-legible
ids (tuomo_suntola, ale_1974_patent) over generated hashes — this data
is meant to be hand-editable.
A person entity may carry an optional portrait — a photo, hotlinked (not
embedded) from its source, so the file stays small and the credit stays
live and checkable. Never add one from an image search or by eyeballing a
"looks about right" photo. The required procedure:
- Find the candidate's Wikidata item (
wbsearchentities, or a web search for"<name>" wikidata). - Verify identity before trusting anything else about that item:
birth_year_verified: this subject's ownevents.jsonalready has abirthevent for the person, and the Wikidata item's P569 matches it. Strongest signal — a coincidence at this specificity is very unlikely.description_verified: no birth event to check against, but the Wikidata description/occupation is specific enough that it couldn't plausibly be a namesake (e.g. "Japanese physicist (1926–2018)" for a Japanese physicist the source discusses in that era — not just "researcher" or "academic", which are too generic to rule out a different person entirely).- Anything weaker (name matches but the description is generic, or contradicts known facts like nationality/era) is not verified — leave the entity without a portrait rather than guessing. A wrong photo is worse than none.
- Only if verified, and the item has a P18 image: resolve the actual file
and its license via the Commons API
(
action=query&prop=imageinfo&iiprop=url|extmetadata), not by guessing a filename. Use the Special:FilePath stable link (https://commons.wikimedia.org/wiki/Special:FilePath/<file>?width=200) asportrait.image_urlso it keeps resolving even if the underlying file is renamed, and record the Commons file page assource_urlso a reader can check the license and original themselves. - Fetching Wikidata/Commons requires a real browser context (fetch calls
from a sandboxed shell are typically blocked) — use whatever browser
tooling is available, not raw
curl.
event_type and relation type are closed enums on purpose — that's what
keeps extraction consistent. To add a value, edit the schema file and add
one line to this README's list rather than letting free-text types
accumulate across subjects. For a new relation type, also add its
(source_types, target_types) entry to RELATION_DIRECTIONS in
scripts/build_site.py — a type missing from that table silently skips
the direction check instead of enforcing it.
For where these vocabularies do (and don't) already exist in authoritative semantic-web ontologies — CIDOC-CRM, schema.org, PROV-O, and others — see Ontology alignment.