Named Graph URI Scheme for CIM Model Ingestion¶
Context and Problem Statement¶
The reactive ingestion pipeline stores every CIM profile, composed base model, SHACL shape set, ontology/vocabulary, and reasoning output as an isolated Named Graph in Apache Jena Fuseki. Each of these graphs goes through multiple lifecycle stages (staged → sealed/locked → quarantined, or composed → versioned) and originates from different logical entities (source system, profile type, model, reasoning run).
Without a consistent, structured identifier scheme, named graph URIs risk becoming ad hoc strings that:
- Cannot be reliably parsed/validated by tooling.
- Conflate mutable lifecycle state with graph identity (forcing renames on every state transition).
- Lose provenance context on failure paths (e.g., a quarantined graph no longer identifiable by its originating system/ingestion run).
- Collide with identifiers used by other RDF tooling or external federated datasets.
How should named graph URIs be structured so that every graph in the system (profile, model, SHACL, vocabulary, reasoning output) is uniquely, deterministically, and human-parseable identified across its entire lifecycle, without relying on mutable or semantically loaded literal names such as staging/production?
Decision Drivers¶
- Referential integrity across lifecycle transitions — a graph's identifier must never need to be renamed or overwritten as it moves from staging to sealed/locked/quarantined; provenance must be traceable from the URI alone.
- Collision avoidance — the scheme must be safe to use even if the triple store is later federated with external CIM/GIS/EMS datasets.
- Human readability / debuggability — operators must be able to recognize a graph's category, source system, and lifecycle stage directly from the Fuseki graph list or SPARQL logs, without a separate lookup table.
- No semantic ambiguity in variable data — instance-specific identifiers (source system, ingestion run, composition run, reasoning run) must be immutable, collision-free values (GUIDs), not human-chosen names that require governance.
- No redundant identifiers — a lifecycle-stage transition must not mint an additional GUID when the parent URN (e.g. the ingestion run's staging URN) is already sufficient to guarantee uniqueness; appending a hash-derived GUID that carries no independent information only adds noise.
- Support for the Composition Engine, SHACL/In-Situ Patching Gate, and Reasoning/Inference write-back described in the ingestion architecture, including quarantine of failed validations and full audit lineage of inferred triples.
Decision Outcome¶
Chosen option: Labeled-segment URN scheme rooted at urn:gmss:graph:, because it satisfies all decision drivers simultaneously: it uses a short, organization-scoped, collision-safe Namespace Identifier (gmss); it encodes graph category and lifecycle stage as inline, self-describing label:value pairs rather than positional-only GUIDs; and it never reuses or mutates a URI across a lifecycle transition — every stage transition appends an additional identifying segment to the full lineage chain, rather than replacing a name (e.g., staging/production) with another.
Within this scheme, two identity patterns are used depending on whether the transition introduces genuinely new information:
- Model/SHACL/Vocabulary/Reasoning categories mint a new GUID at each lifecycle segment (
composition,version,quarantine,reasoning,output), because each of these transitions can occur multiple times for the same parent entity (e.g. multiple composition attempts, multiple locked versions, multiple reasoning runs) and the GUID is what disambiguates them. - Profile ingestion lifecycle segments (
sealed,quarantine,report) are plain literal suffixes with no additional GUID, because they are terminal, at-most-once outcomes of a single staging graph — the staging URN's owningestionGUID already uniquely identifies the run, so appending a further hash-derived GUID would be redundant and would only reduce readability.
URN Grammar¶
Worked Examples¶
| Stage | URN |
|---|---|
| Profile — staging | urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:profile:eq |
| Profile — sealed | urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:profile:eq:sealed |
| Profile — quarantined | urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:profile:eq:quarantine |
| Profile — report | urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:profile:eq:report |
| Profile — staging, consolidated model | urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455 |
| Profile — sealed, consolidated model | urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:sealed |
| Model — composing | urn:gmss:graph:composition:11112222-3333-4444-5555-666677778888:model:6b2e0e4d-1a2b-4c3d-9e8f-abc123456789: |
| Model — locked | urn:gmss:graph:composition:11112222-3333-4444-5555-666677778888:model:6b2e0e4d-1a2b-4c3d-9e8f-abc123456789:version:99990000-aaaa-bbbb-cccc-ddddeeeeffff |
| Model — quarantined | urn:gmss:graph:composition:11112222-3333-4444-5555-666677778888:model:6b2e0e4d-1a2b-4c3d-9e8f-abc123456789:quarantine:7c9e6679-7425-40de-944b-e07fc1f90ae7 |
| SHACL shapes | urn:gmss:graph:shacl:a1b2c3d4-e5f6-4789-90ab-cdef01234567:version:00010002-0003-0004-0005-000600070008 |
| Vocabulary/ontology | urn:gmss:graph:vocab:f0e1d2c3-b4a5-4968-8776-655443322110:version:00000001-0000-4000-8000-000000000001 |
| Reasoning output | urn:gmss:graph:model:6b2e0e4d-1a2b-4c3d-9e8f-abc123456789:composition:11112222-3333-4444-5555-666677778888:version:99990000-aaaa-bbbb-cccc-ddddeeeeffff:reasoning:2f5b8c3a-1234-4e56-8abc-1234567890ab:output:9d8e7f6a-4321-4b56-9cde-abcdef012345 |
Consolidated / Non-Profile Models¶
Some ingested payloads are not a single CIM profile (e.g. a merged/consolidated model spanning multiple profiles in one file). For these, the trailing :profile:{profileType} segment is omitted entirely from the staging URN, and consequently from every derived :sealed, :quarantine, and :report URN as well — there is no profile code to encode. The builder treats an empty/unspecified profile type, or the literal code consolidated, as this case.
Provenance & Retention¶
- Lifecycle state transitions (staging → sealed, composition → locked, any → quarantine) are always additive — a new identifying segment (and, where disambiguation is needed, a new GUID) is appended to the full existing lineage chain; nothing is renamed, overwritten, or stripped of context.
- Cross-URN provenance (e.g., "this sealed profile was sealed from this staging ingestion", "this locked model was locked from this composition run") is additionally recorded as explicit metadata triples (
gmss:sealedFrom,gmss:lockedFrom,gmss:quarantinedFrom, PROV-Oprov:wasGeneratedBy/prov:usedfor reasoning runs) in a side metadata graph, persisted to the KurrentDB event log. - Quarantined and reasoning/inferred graphs are never automatically deleted. Retention is fully operator-controlled; each graph carries a
gmss:retentionPolicy "Manual"metadata flag, and removal only occurs via an explicit operator command (e.g., aDeleteQuarantinedGraphCommand) that performsDROP GRAPH.
Consequences¶
- Good, because every graph's full lineage (source system, ingestion run, profile type, lifecycle stage) is recoverable directly from its URI, without a separate lookup service.
- Good, because lifecycle transitions never require renaming a graph or updating references held by other components (manifests, event logs) — new URNs are minted additively and linked via provenance triples instead.
- Good, because the
gmssNamespace Identifier avoids collisions with generic terms (graph,digital-twin) or other organizations' informal URN usage, while the labeled segments (system:,profile:,ingestion:, etc.) keep the identifiers self-describing without requiring positional memorization. - Good, because profile-lifecycle suffixes (
:sealed,:quarantine,:report) carry no redundant derived GUID — the ingestion run's own GUID already guarantees uniqueness, keeping these URNs shorter and directly predictable from the staging URN. - Good, because reasoning/inference outputs carry full lineage back to the exact model version and composition run that produced them, satisfying audit/traceability requirements for engineering heuristics (e.g., extrapolated zero-sequence impedance).
- Neutral, because URNs grow longer than a purely positional GUID scheme (~110–160 characters for the longest shapes); this remains well within practical limits for SPARQL, Fuseki storage, and RDF serialization.
- Bad, because the grammar requires a dedicated parser/builder component (
NamedGraphUri) to construct and validate these URNs consistently; ad hoc string concatenation elsewhere in the codebase must be avoided to prevent malformed identifiers. - Bad, because quarantine and reasoning-output graphs, having no automatic expiry, require operators to actively manage storage growth over time.
Confirmation¶
Compliance will be validated by:
- A shared
NamedGraphUribuilder/parser class (single source of truth for constructing and validating URNs per this grammar), used by all ingestion, composition, and reasoning components instead of ad hoc string formatting. - Unit tests asserting round-trip parse/format for every category and lifecycle shape defined above, including the consolidated-model (no
profile:segment) variant. - A code review checklist item requiring any new graph-producing component to go through
NamedGraphUrirather than hand-rolled URI strings.
More Information¶
This scheme underpins the ingestion pipeline described in the Base Model Ingestion design (Reactive Ingestion Architecture, Model Composition Engine, and Reasoning/Inference write-back). Related future ADRs may cover: the metadata/provenance graph schema (PROV-O usage), the operator API for quarantine/reasoning-graph deletion, and the system/model registry resolving system/model GUIDs to human-readable display names.