Review phase: surface and cluster untyped entities to suggest ontology extensions #589
kalyp-angel
started this conversation in
Ideas
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Context
The Review tab already has a strong set of queues for graph hygiene and schema consistency: Duplicates, Temporal conflicts, Unconfirmed ("No longer stated"), Low-confidence facts, Axiom violations, and Ontology self-consistency. These cover "is the data clean" and "is the schema internally consistent" very well.
The gap
When extraction encounters a fact the current ontology has no class for, it correctly creates an untyped entity to hold it rather than forcing a bad classification into the nearest existing class, or dropping the fact. That's the right behavior — but there's currently no dedicated place to review these systematically, and no way to see whether the same shape of untyped fact is recurring across multiple documents.
In practice this means a real, recurring gap in the ontology (e.g. several sources all producing an untyped entity with the same relation predicate, is only visible if someone happens to notice it while browsing the graph. One-off untyped entities and a genuine missing-class signal currently look identical.
Proposed feature
An eighth Review queue — something like "Unclassified" — sitting alongside the
existing seven, with two parts:
This feels like the same philosophy already expressed in the Axioms queue's own copy — "a predicate that declares no axioms is never checked" — applied one level up: right now, a class that hasn't been created yet is never surfaced as a pattern, only as scattered individual entities you have to notice by hand.
Why this matters
For anyone doing incremental ontology-building against real documents (rather than starting from a complete schema), this queue would turn "I happened to spot this while looking at the graph" into a repeatable, first-class part of the review workflow —closing the loop between what the extractor is actually finding in the text and what the ontology currently has room for.
Happy to share more detail on the specific workflow that led to this if it's useful.
Curious whether this matches a pattern you're already seeing from other users building ontologies incrementally against real documents, or whether this comes up rarely because most bases start from a more complete schema — genuinely not sure how common this scenario is outside our own case, and that would shape how worth prioritizing it is.
All reactions