Skip to content

Named Graph URI Scheme for CIM Model Ingestion

Context and Problem Statement

The reactive ingestion pipeline stores every CIM profile, composed base model, SHACL shape set, ontology/vocabulary, and reasoning output as an isolated Named Graph in Apache Jena Fuseki. Each of these graphs goes through multiple lifecycle stages (staged → sealed/locked → quarantined, or composed → versioned) and originates from different logical entities (source system, profile type, model, reasoning run).

Without a consistent, structured identifier scheme, named graph URIs risk becoming ad hoc strings that:

  • Cannot be reliably parsed/validated by tooling.
  • Conflate mutable lifecycle state with graph identity (forcing renames on every state transition).
  • Lose provenance context on failure paths (e.g., a quarantined graph no longer identifiable by its originating system/ingestion run).
  • Collide with identifiers used by other RDF tooling or external federated datasets.

How should named graph URIs be structured so that every graph in the system (profile, model, SHACL, vocabulary, reasoning output) is uniquely, deterministically, and human-parseable identified across its entire lifecycle, without relying on mutable or semantically loaded literal names such as staging/production?

Decision Drivers

  • Referential integrity across lifecycle transitions — a graph's identifier must never need to be renamed or overwritten as it moves from staging to sealed/locked/quarantined; provenance must be traceable from the URI alone.
  • Collision avoidance — the scheme must be safe to use even if the triple store is later federated with external CIM/GIS/EMS datasets.
  • Human readability / debuggability — operators must be able to recognize a graph's category, source system, and lifecycle stage directly from the Fuseki graph list or SPARQL logs, without a separate lookup table.
  • No semantic ambiguity in variable data — instance-specific identifiers (source system, ingestion run, composition run, reasoning run) must be immutable, collision-free values (GUIDs), not human-chosen names that require governance.
  • No redundant identifiers — a lifecycle-stage transition must not mint an additional GUID when the parent URN (e.g. the ingestion run's staging URN) is already sufficient to guarantee uniqueness; appending a hash-derived GUID that carries no independent information only adds noise.
  • Support for the Composition Engine, SHACL/In-Situ Patching Gate, and Reasoning/Inference write-back described in the ingestion architecture, including quarantine of failed validations and full audit lineage of inferred triples.

Decision Outcome

Chosen option: Labeled-segment URN scheme rooted at urn:gmss:graph:, because it satisfies all decision drivers simultaneously: it uses a short, organization-scoped, collision-safe Namespace Identifier (gmss); it encodes graph category and lifecycle stage as inline, self-describing label:value pairs rather than positional-only GUIDs; and it never reuses or mutates a URI across a lifecycle transition — every stage transition appends an additional identifying segment to the full lineage chain, rather than replacing a name (e.g., staging/production) with another.

Within this scheme, two identity patterns are used depending on whether the transition introduces genuinely new information:

  • Model/SHACL/Vocabulary/Reasoning categories mint a new GUID at each lifecycle segment (composition, version, quarantine, reasoning, output), because each of these transitions can occur multiple times for the same parent entity (e.g. multiple composition attempts, multiple locked versions, multiple reasoning runs) and the GUID is what disambiguates them.
  • Profile ingestion lifecycle segments (sealed, quarantine, report) are plain literal suffixes with no additional GUID, because they are terminal, at-most-once outcomes of a single staging graph — the staging URN's own ingestion GUID already uniquely identifies the run, so appending a further hash-derived GUID would be redundant and would only reduce readability.

URN Grammar

Worked Examples

Stage URN
Profile — staging urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:profile:eq
Profile — sealed urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:profile:eq:sealed
Profile — quarantined urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:profile:eq:quarantine
Profile — report urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:profile:eq:report
Profile — staging, consolidated model urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455
Profile — sealed, consolidated model urn:gmss:graph:ingestion:8f4a12a9-2e02-4c28-9c31-77a0e1f9d001:system:3f9a1b2c-91de-4a55-9c2e-001122334455:sealed
Model — composing urn:gmss:graph:composition:11112222-3333-4444-5555-666677778888:model:6b2e0e4d-1a2b-4c3d-9e8f-abc123456789:
Model — locked urn:gmss:graph:composition:11112222-3333-4444-5555-666677778888:model:6b2e0e4d-1a2b-4c3d-9e8f-abc123456789:version:99990000-aaaa-bbbb-cccc-ddddeeeeffff
Model — quarantined urn:gmss:graph:composition:11112222-3333-4444-5555-666677778888:model:6b2e0e4d-1a2b-4c3d-9e8f-abc123456789:quarantine:7c9e6679-7425-40de-944b-e07fc1f90ae7
SHACL shapes urn:gmss:graph:shacl:a1b2c3d4-e5f6-4789-90ab-cdef01234567:version:00010002-0003-0004-0005-000600070008
Vocabulary/ontology urn:gmss:graph:vocab:f0e1d2c3-b4a5-4968-8776-655443322110:version:00000001-0000-4000-8000-000000000001
Reasoning output urn:gmss:graph:model:6b2e0e4d-1a2b-4c3d-9e8f-abc123456789:composition:11112222-3333-4444-5555-666677778888:version:99990000-aaaa-bbbb-cccc-ddddeeeeffff:reasoning:2f5b8c3a-1234-4e56-8abc-1234567890ab:output:9d8e7f6a-4321-4b56-9cde-abcdef012345

Consolidated / Non-Profile Models

Some ingested payloads are not a single CIM profile (e.g. a merged/consolidated model spanning multiple profiles in one file). For these, the trailing :profile:{profileType} segment is omitted entirely from the staging URN, and consequently from every derived :sealed, :quarantine, and :report URN as well — there is no profile code to encode. The builder treats an empty/unspecified profile type, or the literal code consolidated, as this case.

Provenance & Retention

  • Lifecycle state transitions (staging → sealed, composition → locked, any → quarantine) are always additive — a new identifying segment (and, where disambiguation is needed, a new GUID) is appended to the full existing lineage chain; nothing is renamed, overwritten, or stripped of context.
  • Cross-URN provenance (e.g., "this sealed profile was sealed from this staging ingestion", "this locked model was locked from this composition run") is additionally recorded as explicit metadata triples (gmss:sealedFrom, gmss:lockedFrom, gmss:quarantinedFrom, PROV-O prov:wasGeneratedBy / prov:used for reasoning runs) in a side metadata graph, persisted to the KurrentDB event log.
  • Quarantined and reasoning/inferred graphs are never automatically deleted. Retention is fully operator-controlled; each graph carries a gmss:retentionPolicy "Manual" metadata flag, and removal only occurs via an explicit operator command (e.g., a DeleteQuarantinedGraphCommand) that performs DROP GRAPH.

Consequences

  • Good, because every graph's full lineage (source system, ingestion run, profile type, lifecycle stage) is recoverable directly from its URI, without a separate lookup service.
  • Good, because lifecycle transitions never require renaming a graph or updating references held by other components (manifests, event logs) — new URNs are minted additively and linked via provenance triples instead.
  • Good, because the gmss Namespace Identifier avoids collisions with generic terms (graph, digital-twin) or other organizations' informal URN usage, while the labeled segments (system:, profile:, ingestion:, etc.) keep the identifiers self-describing without requiring positional memorization.
  • Good, because profile-lifecycle suffixes (:sealed, :quarantine, :report) carry no redundant derived GUID — the ingestion run's own GUID already guarantees uniqueness, keeping these URNs shorter and directly predictable from the staging URN.
  • Good, because reasoning/inference outputs carry full lineage back to the exact model version and composition run that produced them, satisfying audit/traceability requirements for engineering heuristics (e.g., extrapolated zero-sequence impedance).
  • Neutral, because URNs grow longer than a purely positional GUID scheme (~110–160 characters for the longest shapes); this remains well within practical limits for SPARQL, Fuseki storage, and RDF serialization.
  • Bad, because the grammar requires a dedicated parser/builder component (NamedGraphUri) to construct and validate these URNs consistently; ad hoc string concatenation elsewhere in the codebase must be avoided to prevent malformed identifiers.
  • Bad, because quarantine and reasoning-output graphs, having no automatic expiry, require operators to actively manage storage growth over time.

Confirmation

Compliance will be validated by:

  • A shared NamedGraphUri builder/parser class (single source of truth for constructing and validating URNs per this grammar), used by all ingestion, composition, and reasoning components instead of ad hoc string formatting.
  • Unit tests asserting round-trip parse/format for every category and lifecycle shape defined above, including the consolidated-model (no profile: segment) variant.
  • A code review checklist item requiring any new graph-producing component to go through NamedGraphUri rather than hand-rolled URI strings.

More Information

This scheme underpins the ingestion pipeline described in the Base Model Ingestion design (Reactive Ingestion Architecture, Model Composition Engine, and Reasoning/Inference write-back). Related future ADRs may cover: the metadata/provenance graph schema (PROV-O usage), the operator API for quarantine/reasoning-graph deletion, and the system/model registry resolving system/model GUIDs to human-readable display names.