How Semantica's Decision Intelligence Layer Works: Semantic-Structural Hybrid Architecture
Semantica’s decision intelligence layer combines semantic embeddings from NLP with structural embeddings from knowledge graphs to record, search, and explain business decisions through a three-component architecture comprising DecisionContext, DecisionEmbeddingPipeline, and decision vector methods.
The Semantica framework provides an open-source decision-intelligence layer that transforms raw business decisions into searchable, explorable knowledge. Located in the semantica-agi/semantica repository, this layer enables organizations to capture decision context and retrieve relevant precedents using hybrid similarity search. This article examines the source code implementation to explain how the system processes, embeds, and retrieves decision data.
Core Architecture of the Decision Intelligence Layer
The decision intelligence layer consists of three tightly-coupled components that orchestrate the flow from raw input to retrievable knowledge.
- DecisionContext (
semantica/context/decision_context.py): The high-level API that records decisions, executes hybrid precedent searches, generates explanations, and aggregates statistics. - DecisionEmbeddingPipeline (
semantica/vector_store/decision_embedding_pipeline.py): Generates dual embeddings—semantic vectors from text models and structural vectors from knowledge graphs using Node2Vec and optional KG algorithms. - Decision Vector Methods (
semantica/vector_store/decision_vector_methods.py): Convenience wrappers includingquick_decision,find_precedents, andexplainthat expose pipeline functionality as one-line operations for developers and CLI tools.
Recording Decisions: The Seven-Stage Pipeline
When DecisionContext.record_decision() is invoked (lines 60-87 in decision_context.py), the system executes a seven-stage validation and embedding pipeline defined in DecisionEmbeddingPipeline.process_decision():
-
Data Validation:
_validate_decision_dataensures the mandatoryscenariofield exists and populates defaults for optional fields. -
Semantic Embedding Generation:
_generate_semantic_embeddingconcatenates thescenario,reasoning,outcome, andcategoryfields into a single text string, then callsvector_store.embed(text). If the embedding call fails, the system falls back to a random vector. -
Structural Embedding Generation: When a graph store is present,
_generate_structural_embeddinginstantiates aNodeEmbedder(Node2Vec implementation fromsemantica/kg/node_embeddings.py) to compute embeddings for decision entities or categories. Withuse_graph_featuresenabled, additional KG algorithms—path-finders, community detectors, and centrality calculators—enrich the base embeddings. -
Vector Combination:
_create_combined_embeddingblends semantic and structural vectors using configurable weights with defaults ofsemantic_weight=0.7andstructural_weight=0.3. -
Metadata Enrichment:
_enrich_metadatainjects pipeline version, timestamps, and current weight settings into the decision record. -
Storage:
_store_embeddingswrites the semantic vector (and optionally the structural vector) to the vector store, returning a unique vector ID. -
Return: The vector ID propagates back through the context layer to the caller.
Hybrid Precedent Search
The DecisionContext.find_similar_decisions() method delegates retrieval to the ContextRetriever class, which implements hybrid search through the following sequence.
- Query Processing: The query decision is converted to a temporary embedding via
process_decision(..., store_embeddings=False)without persisting to storage. - Candidate Retrieval:
_get_candidate_embeddingsextracts a large pool of candidate vectors from the vector store, optionally filtered, including any stored structural embeddings. - Hybrid Similarity Calculation:
HybridSimilarityCalculator.find_most_similar_decisionscombines semantic and structural similarities using current weights. If the backend returns scores instead of raw vectors, the method patches semantic similarity with backend scores (see lines 44-52 ofdecision_embedding_pipeline.py). - Result Ranking: Combined similarity scores are sorted, and the top-k decisions are returned to the caller.
Developers can disable hybrid search (semantic-only) or force structural-only search via the use_hybrid_search parameter.
Decision Explanation and Transparency
The explain_decision method (lines 65-78 in decision_context.py) provides interpretability through a hierarchical fallback strategy.
First, the method attempts to use the vector store's native explain_decision capability. If unavailable, it constructs a structured explanation from metadata including scenario, reasoning, outcome, confidence scores, and current semantic/structural weight values. When include_paths=True, the explanation incorporates similar decisions retrieved via find_similar_decisions to illustrate reasoning paths and precedent relationships.
Configuration and Batch Processing
Dynamic Weight Adjustment
Both DecisionContext and DecisionEmbeddingPipeline expose update_similarity_weights to modify search behavior at runtime.
def update_weights(self, semantic_weight: float, structural_weight: float) -> None:
self.semantic_weight = semantic_weight
self.structural_weight = structural_weight
self.hybrid_calculator.update_weights(semantic_weight, structural_weight)
This method, located at lines 302-308 of decision_embedding_pipeline.py, immediately updates the HybridSimilarityCalculator instance, affecting all subsequent searches without requiring pipeline reinitialization.
Batch Processing
For high-throughput scenarios, batch_decisions (or process_decision_batch) pre-generates structural embeddings for all unique entities to avoid redundant computation. The method processes decisions in configurable batch sizes (default 32), tracking progress via ProgressTracker (see lines 42-80 of the pipeline implementation).
Statistics and Monitoring
The get_statistics method (aliased as stats) reports operational metrics including total decisions processed, current weight configurations, graph store availability, cached structural embedding counts, and metadata distributions across categories and outcomes.
Practical Implementation Examples
The following examples demonstrate common operations using the decision intelligence layer.
from semantica.context import DecisionContext
# Initialize context with vector and graph stores
ctx = DecisionContext(vector_store=vs, graph_store=gs)
# Record a decision with full context
decision_id = ctx.record_decision(
scenario="Credit limit increase for VIP client",
reasoning="Excellent repayment history, low risk score",
outcome="approved",
confidence=0.92,
entities=["client_123", "credit_account"],
category="finance"
)
# Retrieve hybrid precedent decisions
precedents = ctx.find_similar_decisions(
scenario="Credit limit increase for VIP client",
limit=5,
use_hybrid_search=True
)
# Generate explanation with reasoning paths
explanation = ctx.explain_decision(decision_id, include_paths=True)
# Batch process multiple decisions
batch = [
{"scenario": "Loan approval", "outcome": "rejected", "entities": ["loan_001"]},
{"scenario": "Insurance claim", "outcome": "approved", "entities": ["claim_42"]},
]
results = ctx.process_decision_batch(batch, batch_size=2)
# Adjust search weights dynamically
ctx.update_similarity_weights(semantic_weight=0.6, structural_weight=0.4)
# Retrieve operational statistics
stats = ctx.get_statistics()
Summary
- Three-component architecture: The decision intelligence layer combines
DecisionContext(orchestration),DecisionEmbeddingPipeline(dual embedding generation), and vector methods (convenience API). - Dual embedding strategy: Semantic embeddings capture textual meaning via NLP models, while structural embeddings encode relational context via Node2Vec and knowledge graph algorithms.
- Configurable hybrid search: The system blends embedding types using adjustable weights (default 70% semantic, 30% structural) with runtime modification capabilities.
- Full provenance tracking: Decision recording includes validation, enrichment, and storage phases that maintain traceability from raw input to vector ID.
- Explainable AI support: Native explanation methods provide transparency through metadata assembly and precedent path visualization.
Frequently Asked Questions
How does Semantica combine semantic and structural embeddings?
The system generates semantic embeddings from text fields (scenario, reasoning, outcome, category) using the vector store's embedding model, and structural embeddings from knowledge graph entities using Node2Vec. The _create_combined_embedding method in DecisionEmbeddingPipeline blends these vectors using semantic_weight and structural_weight parameters, defaulting to 0.7 and 0.3 respectively. The HybridSimilarityCalculator class in semantica/vector_store/hybrid_similarity.py performs the weighted combination during search operations.
What happens if the embedding service fails during decision recording?
The pipeline includes resilience mechanisms in _generate_semantic_embedding. If the call to vector_store.embed(text) raises an exception, the method falls back to generating a random vector, ensuring the decision recording process continues without interruption. This prevents data loss during temporary vector store outages while maintaining system availability.
Can I use the decision intelligence layer without a knowledge graph?
Yes. While the system supports enhanced structural embeddings when graph_store is provided to DecisionContext, it operates in semantic-only mode if the graph store is absent. The use_graph_features flag and presence checks in the pipeline ensure that structural embedding generation is skipped when no graph is available, with the hybrid calculator automatically adjusting to use only semantic similarity scores.
How does batch processing improve performance?
The process_decision_batch method optimizes throughput by pre-generating structural embeddings for all unique entities across the batch before processing individual decisions. This eliminates redundant Node2Vec computations when multiple decisions reference the same entities. Additionally, configurable batch sizing (default 32) and ProgressTracker integration enable efficient memory management and progress monitoring during high-volume ingestion scenarios.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →