What Is a Context Graph in Semantica? Architecture and Implementation Guide
A Context Graph in Semantica is an in-memory knowledge graph implementation that stores entities and relationships with temporal validity windows, decision provenance, and advanced analytics capabilities.
The Context Graph serves as the central data structure in the semantica-agi/semantica repository, enabling AI agents to build knowledge representations from documents and conversations. This thread-safe graph store not only manages static facts but tracks when information is valid, records decision provenance, and supports advanced graph analytics such as centrality and community detection.
Core Architecture and Data Model
Nodes, Edges, and Temporal Metadata
At the foundation of the Context Graph are two dataclasses defined in semantica/context/context_graph.py (lines 18-29): ContextNode and ContextEdge. These structures hold unique identifiers, types, content payloads, and rich metadata. Crucially, both support optional valid_from and valid_until timestamps, enabling time-aware knowledge representation where facts can expire or become active at specific moments.
The temporal fields are normalized to timezone-naive UTC ISO-8601 strings through internal helper methods _parse_iso_dt and _normalize_temporal_input (lines 23-46). Both dataclasses expose an is_active property that evaluates whether the current time falls within the validity window, allowing queries to filter out retracted or future-dated information automatically.
Thread-Safe Mutations
All graph mutations are protected by a re-entrant lock instantiated at line 100: self._lock = threading.RLock(). This ensures that concurrent operations—such as adding nodes during decision recording or updating edges during batch ingestion—maintain graph integrity without race conditions. The lock guards every CRUD operation, including add_node, add_edge, add_nodes, and add_edges (lines 124-167).
Advanced Capabilities
Temporal Validity and Retraction
Unlike static knowledge graphs, the Context Graph in Semantica treats time as a first-class citizen. The _default_edge_id method (lines 27-40) generates deterministic UUID5 identifiers from edge payloads, while _resolve_edge_identity handles deduplication. When combined with temporal metadata, these mechanisms enable sophisticated patterns such as automatic edge retraction or historical versioning without deleting past state.
Decision Provenance Tracking
The graph integrates decision management through methods like record_decision, find_precedents, analyze_decision_influence, and get_decision_insights (referenced in lines 14-22). These APIs weave decision provenance directly into the graph structure, storing not just what was decided but why, with full entity attribution and confidence scoring. This supports causal analysis and policy enforcement by tracing how specific entities influenced past outcomes.
Edge Deduplication and Identity
To prevent duplicate relationships, edge identity is computed deterministically using UUID5 hashes generated by _default_edge_id from the source node, target node, relationship type, and payload content. Developers can override this behavior through _resolve_edge_identity, but the default implementation ensures that identical relationships collapse to a single edge even when added multiple times from different ingestion pipelines.
Graph Analytics and Querying
Traversal and Search
The Context Graph provides BFS-based neighbor discovery through get_neighbors and distance-aware retrieval via get_neighbor_distances (lines 188-214). These methods support hop-limited traversal with configurable minimum weights and can include distance metadata showing confidence decay across multi-hop relationships. For content-based discovery, the query method performs simple keyword search across node content and metadata.
Optional KG Components
When initialized with advanced_analytics=True, the graph instantiates additional computational components (lines 140-156) from the optional semantica/kg package. These include algorithms for centrality calculation, community detection, node embeddings, and path-finding. These capabilities are loaded conditionally, allowing the core graph to remain lightweight while supporting sophisticated AI reasoning when needed.
Practical Implementation Examples
Initializing and Populating the Graph
from semantica.context import ContextGraph
# Initialise with advanced analytics enabled (requires optional KG deps)
graph = ContextGraph(advanced_analytics=True)
# Add some entities
graph.add_node("Python", "language", content="Python programming language")
graph.add_node("Programming", "concept")
graph.add_edge("Python", "Programming", "related_to")
Temporal Validity and Campaign Management
# Add a node that is only valid for 2025
graph.add_node(
"Promo2025",
"campaign",
content="2025 Summer Promotion",
valid_from="2025-06-01T00:00:00",
valid_until="2025-09-30T23:59:59",
)
# Query only currently active nodes
active_campaigns = [n for n in graph.nodes.values() if n.is_active]
Multi-Hop Traversal with Metadata
neighbors = graph.get_neighbors(
"Python",
hops=2,
min_weight=0.5,
include_distance_metadata=True,
)
for n in neighbors:
print(f"{n['id']} ({n['type']}) – hop {n['hop']}, decay {n['confidence_decay']:.2f}")
Keyword Search and Export
# Simple content search
results = graph.query("programming language", limit=5)
# Export to version-controlled markdown
md = graph.export_node_markdown("Python")
print(md) # markdown with front-matter
Key Source Files
The Context Graph implementation spans several files in the Semantica repository:
semantica/context/context_graph.py– Core implementation containingContextNode,ContextEdge, CRUD operations, traversal methods, and decision-tracking integration (lines 14-214+).semantica/utils/helpers.py– Utility functions such asclassify_path_distanceused for distance metadata calculations.semantica/kg/– Optional knowledge-graph algorithms including centrality, community detection, and embeddings that integrate whenadvanced_analyticsis enabled.semantica/visualization/visualization_provenance.py– Utilities for rendering graph provenance and decision flows.examples/capability_gap_context_graphs_example.py– End-to-end example demonstrating capability-gap analysis using the Context Graph.
Summary
- The Context Graph is an in-memory, thread-safe knowledge graph using
ContextNodeandContextEdgedataclasses with built-in temporal validity. - Located in
semantica/context/context_graph.py, it provides CRUD operations guarded bythreading.RLock()for concurrent access. - Temporal features use ISO-8601 normalization via
_parse_iso_dtand_normalize_temporal_input, withis_activeproperties filtering time-bound entities. - Decision provenance is native to the graph through
record_decisionand related methods, enabling causal analysis and precedent search. - Edge deduplication uses deterministic UUID5 generation in
_default_edge_idto prevent duplicate relationships. - Optional advanced analytics (centrality, embeddings, community detection) load when
advanced_analytics=Trueis passed to the constructor.
Frequently Asked Questions
How does temporal validity work in the Context Graph?
Temporal validity is implemented through valid_from and valid_until timestamp fields on both nodes and edges. These are normalized to timezone-naive UTC ISO-8601 strings using _normalize_temporal_input and _parse_iso_dt. The is_active property checks whether the current time falls within these bounds, allowing queries to automatically exclude expired or future-dated information without deleting historical data.
What makes the Context Graph thread-safe?
Thread safety is enforced by a re-entrant lock (self._lock = threading.RLock()) initialized at line 100 of context_graph.py. This lock guards all mutation operations including add_node, add_edge, and their batched variants (add_nodes, add_edges), ensuring that concurrent modifications from multiple threads or async tasks cannot corrupt the internal node and edge indexes.
How does the Context Graph prevent duplicate edges?
Edge identity is computed deterministically using UUID5 hashes generated by the _default_edge_id method (lines 27-40), which derives the edge ID from the source node, target node, relationship type, and content payload. The _resolve_edge_identity helper applies this logic during edge creation, ensuring that identical relationships collapse to a single edge even when ingested from multiple sources.
Can the Context Graph track why decisions were made?
Yes, the graph includes native decision provenance APIs including record_decision, find_precedents, and analyze_decision_influence. These methods store decision metadata—such as reasoning, confidence scores, and affected entities—directly within the graph structure. This enables backward-looking causal analysis, precedent-based reasoning, and policy enforcement by tracing the influence of specific entities on historical outcomes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →