How Semantica's ContextGraph Performs Causal Chain Analysis: BFS Traversal and Edge Semantics
Semantica's ContextGraph performs causal chain analysis by executing a breadth-first search (BFS) in get_causal_chain, filtering edges against the _CAUSAL_TRAVERSAL_TYPES set, and returning Decision dataclasses annotated with causal_distance metadata based on traversal direction and configurable depth limits.
Understanding how Semantica's ContextGraph performs causal chain analysis requires examining the graph traversal implementation in the semantica-agi/semantica repository. The system models AI decisions as nodes in an in-memory graph, where causal relationships like CAUSED and INFLUENCED form directed edges between decision points. This architecture enables deterministic auditing of decision pipelines by tracing upstream dependencies or downstream consequences through configurable traversal depths.
Core Architecture of Causal Chain Analysis
Causal Edge Typing in ContextGraph
The causal chain analysis relies on a specific subset of edge types defined in the _CAUSAL_TRAVERSAL_TYPES collection within semantica/context/context_graph.py. This set includes relationship types such as CAUSED and INFLUENCED that define directional dependencies between decision nodes. When the graph traversal algorithm evaluates connections, it exclusively examines edges whose edge_type belongs to this causal set, ignoring non-causal relationships.
The Decision Dataclass Model
Each node in the causal chain transforms into a Decision dataclass defined in semantica/context/decision_models.py (lines 87-98). This structure preserves the original node properties while augmenting them with traversal-specific metadata. During causal chain analysis, the system populates the metadata dictionary with a causal_distance field representing the node's BFS depth from the origin decision.
The get_causal_chain Traversal Algorithm
Implemented in semantica/context/context_graph.py (lines 3918-4000), the get_causal_chain method executes a deterministic graph traversal to extract decision-centric views of causal influences. The algorithm accepts three primary parameters: decision_id (the origin node UUID), direction (either "upstream" or "downstream"), and max_depth (an integer limiting traversal hops).
Input Validation and Directionality
The method first validates that the direction parameter contains either "upstream" or "downstream", raising a ValueError for invalid inputs. This parameter determines whether the traversal follows causal edges toward predecessors or successors.
- Upstream traversal: Examines incoming edges where the current node serves as the target of a causal relationship, identifying root causes.
- Downstream traversal: Examines outgoing edges where the current node serves as the source of a causal relationship, identifying consequences.
Breadth-First Search with Cycle Detection
The algorithm initializes a queue containing tuples of (node_id, depth), starting with the supplied decision_id at depth 0. A visited set tracks processed nodes to prevent infinite loops in cyclic graphs. The traversal continues until the queue empties or reaches the max_depth limit, ensuring bounded execution time regardless of graph topology.
Edge Filtering and Node Expansion
During each iteration, the algorithm dequeues a node and examines adjacent edges based on the traversal direction. It filters these edges against _CAUSAL_TRAVERSAL_TYPES to ensure only causal relationships contribute to the chain. Valid neighboring nodes undergo type validation—the system only includes nodes where node_type equals "decision", instantiating each as a fully populated Decision object from the decision models module.
Depth-Aware Result Ordering
Each discovered decision receives causal_distance metadata equal to its BFS depth. The final sorting logic varies by direction to present results in logical causal order:
- Upstream chains: Return results with the most distant decisions first (reverse depth order), showing root causes before immediate predecessors.
- Downstream chains: Return results with closest decisions first (natural depth order), showing immediate effects before distant consequences.
Practical Causal Chain Query Examples
The following examples demonstrate causal chain retrieval using the Semantica API.
Retrieve an upstream causal chain with default depth:
from semantica.context import ContextGraph
graph = ContextGraph()
# …populate graph with decisions and causal edges…
chain = graph.get_causal_chain(decision_id="dec_123", direction="upstream")
for d in chain:
print(d.decision_id, d.metadata["causal_distance"])
Retrieve a downstream chain limited to three hops:
downstream = graph.get_causal_chain(
decision_id="dec_123",
direction="downstream",
max_depth=3,
)
# `downstream` now contains at most three decision hops that were caused by `dec_123`
Process chain results in analytics routines:
def summarize_chain(chain):
return {
"total_steps": len(chain),
"ids": [d.decision_id for d in chain],
"max_distance": max(d.metadata["causal_distance"] for d in chain) if chain else 0,
}
summary = summarize_chain(chain)
print(summary)
Summary
- Semantica's ContextGraph stores decisions as nodes and causal relationships (like
CAUSED,INFLUENCED) as typed edges in_CAUSAL_TRAVERSAL_TYPES. - The
get_causal_chainmethod insemantica/context/context_graph.pyimplements BFS traversal with configurablemax_depthand directionality (upstreamordownstream). - The algorithm uses a visited set for cycle detection and filters nodes by
node_type == "decision"before instantiatingDecisiondataclasses fromsemantica/context/decision_models.py. - Results include
causal_distancemetadata and are sorted by depth: upstream chains show most distant first, downstream chains show closest first. - The system raises ValueError for invalid direction parameters, enforcing strict API contracts.
Frequently Asked Questions
What edge types does ContextGraph consider causal?
According to the semantica source code, causal edges belong to the _CAUSAL_TRAVERSAL_TYPES set, which includes types such as CAUSED and INFLUENCED. The get_causal_chain method exclusively traverses edges matching these types, ignoring other relationship categories during causal analysis.
How does get_causal_chain handle cyclic causal relationships?
The implementation prevents infinite loops by maintaining a visited set that tracks processed node IDs. Before enqueuing any neighbor, the algorithm checks against this set, ensuring each decision appears only once in the causal chain regardless of cyclic dependencies in the underlying graph.
What is the difference between upstream and downstream causal chain analysis?
Upstream analysis traces dependencies backward by following edges where the current node is the target, identifying root causes and influencing decisions. Downstream analysis traces consequences forward by following edges where the current node is the source, identifying effects and resulting decisions. The return order also differs: upstream sorts by reverse depth (most distant first), while downstream uses natural depth order.
What is the default maximum depth for causal chain traversal?
While the source code accepts a configurable max_depth parameter, common usage patterns in the semantica repository indicate a default effective depth of 10 hops when not explicitly specified. Developers can override this by passing explicit max_depth values to limit traversal scope and improve query performance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →