Semantica Data Processing Pipeline: 8 Distinct Stages Explained

Semantica's decision-embedding pipeline processes raw decision records through eight distinct stages—validation, semantic embedding generation, structural embedding generation, vector combination, metadata enrichment, storage, batch processing, and similarity search—to produce enriched hybrid vector representations.

The semantica-agi/semantica repository implements a hybrid architecture that combines language-based semantic analysis with graph-based structural context. Understanding the distinct stages in Semantica's data processing pipeline enables developers to optimize ingestion workflows and leverage advanced knowledge-graph features for decision tracking.

Stage 1: Validation and Normalization

The pipeline begins by enforcing data integrity through the _validate_decision_data method in semantica/vector_store/decision_embedding_pipeline.py. This stage checks for required fields such as scenario and applies sensible defaults for optional fields including outcome, reasoning, and confidence.

Input records that lack mandatory fields are rejected at this boundary, ensuring downstream embedding generation receives consistent, well-formed decision objects.

Stage 2: Semantic Embedding Generation

Once validated, decision text is converted into dense vectors via _generate_semantic_embedding. The implementation concatenates relevant text fields—scenario, reasoning, outcome, and category—into a single document representation.

The configured vector_store.embed method generates the semantic vector. If the primary embedding service fails, the pipeline automatically falls back to a random vector, ensuring system continuity at the cost of semantic accuracy.

Stage 3: Structural Embedding Generation (Optional)

When a graph store is configured and use_graph_features=True, the pipeline invokes _generate_structural_embedding to capture relational context. This stage extracts entities from the decision and utilizes NodeEmbedder.compute_embeddings (implemented in semantica/kg/node_embeddings.py) to generate Node2Vec-based representations.

Advanced knowledge-graph algorithms enhance these embeddings:

Stage 4: Combination of Semantic and Structural Vectors

The _create_combined_embedding method merges dual-vector representations into a single hybrid embedding. This stage first resizes vectors to identical dimensionalities if they differ, then applies configurable weights (semantic_weight and structural_weight) to balance linguistic meaning against graph topology.

The default weighting favors semantic content, but users can adjust the ratio to prioritize relational structure for specific use cases.

Stage 5: Metadata Enrichment

Before persistence, _enrich_metadata attaches provenance information to each decision record. Captured metadata includes:

  • Pipeline version identifier
  • Generation timestamps
  • Active weight settings (semantic_weight, structural_weight)
  • Boolean flags indicating structural embedding inclusion

This metadata enables audit trails and facilitates A/B testing of different pipeline configurations.

Stage 6: Storage

The _store_embeddings method persists finalized vectors to the configured vector store. Both the combined hybrid vector and the raw semantic vector are stored alongside enriched metadata, enabling retrieval strategies that query either representation independently.

Storage atomicity ensures that partial failures do not leave orphaned metadata or unreferenced vector entries.

Stage 7: Batch Processing (Optional)

For high-volume ingestion, _process_decision_batch optimizes throughput by pre-generating structural embeddings once per unique entity, then reusing them across the decision batch. This avoids redundant graph computations when processing thousands of related decisions.

Progress tracking is handled by ProgressTracker (from semantica/utils/progress_tracker.py), which reports completion percentages and estimated time remaining for long-running batch jobs.

Stage 8: Similarity Search (Optional)

The pipeline supports hybrid retrieval via _find_similar_decisions. Query decisions traverse the same eight stages to generate comparable vector representations. The system then retrieves candidate embeddings from the vector store and computes hybrid similarity scores combining both semantic proximity and graph structural alignment.

This enables precedent retrieval that respects not just textual similarity but also organizational or causal relationships encoded in the knowledge graph.

Implementation Example

The following code demonstrates the complete pipeline workflow:


# Initialise the pipeline (semantic vector store + optional graph store)

pipeline = DecisionEmbeddingPipeline(
    vector_store=my_vector_store,
    graph_store=my_graph_store,          # omit to disable structural embeddings

    use_graph_features=True,             # enable KG‑enhanced embeddings

    semantic_weight=0.7,
    structural_weight=0.3,
)

# 1️⃣ Process a single decision

result = pipeline.process_decision({
    "scenario": "Increase credit limit",
    "reasoning": "Customer has good payment history",
    "outcome": "approved",
    "entities": ["Customer", "CreditAccount"],
    "category": "Finance",
})
print(result["combined_embedding"].shape)   # → (384,)

# 2️⃣ Process a batch of decisions

batch = [
    {"scenario": "Add new product line", "entities": ["Product"], "category": "Strategy"},
    {"scenario": "Close under‑performing store", "entities": ["Store"], "category": "Operations"},
]
batch_results = pipeline.process_decision_batch(batch, batch_size=2)

# 3️⃣ Find similar past decisions (hybrid search)

similar = pipeline.find_similar_decisions(
    query_decision={"scenario": "Raise credit limit", "entities": ["Customer"]},
    limit=5,
)
for hit in similar:
    print(hit["metadata"]["scenario"], hit["similarity"])

Summary

  • Validation and Normalization ensures data integrity before processing via _validate_decision_data.
  • Semantic Embedding Generation converts text fields into dense vectors using vector_store.embed.
  • Structural Embedding Generation optionally extracts graph features via NodeEmbedder.compute_embeddings.
  • Vector Combination merges representations using configurable weights in _create_combined_embedding.
  • Metadata Enrichment adds timestamps and configuration details through _enrich_metadata.
  • Storage persists hybrid vectors via _store_embeddings.
  • Batch Processing optimizes throughput for large datasets using _process_decision_batch.
  • Similarity Search enables hybrid retrieval matching both text and graph structure via _find_similar_decisions.

Frequently Asked Questions

What is the difference between semantic and structural embeddings in Semantica?

Semantic embeddings capture linguistic meaning by encoding concatenated text fields (scenario, reasoning, outcome) through standard vector store models. Structural embeddings encode relational context extracted from knowledge graphs using Node2Vec algorithms and graph metrics like centrality and community detection. The pipeline combines both to create hybrid representations that understand both what was decided and how it relates to other decisions.

How does Semantica handle invalid decision data during processing?

The _validate_decision_data method in semantica/vector_store/decision_embedding_pipeline.py acts as a strict gatekeeper. Records missing required fields such as scenario are rejected immediately, while optional fields receive default values. This prevents malformed data from corrupting downstream embedding generation or storage operations.

Can I use Semantica's pipeline without a graph store?

Yes. The graph store is optional. If graph_store is omitted or use_graph_features=False, the pipeline skips structural embedding generation (Stage 3) and relies solely on semantic vectors. In this mode, _create_combined_embedding simply passes through the semantic representation, and similarity search operates on text-based vectors only.

What happens if the vector store fails to generate an embedding?

The _generate_semantic_embedding method includes a fault-tolerance mechanism. If vector_store.embed raises an exception or returns an error, the pipeline generates a random fallback vector of the expected dimensionality. This ensures system availability, though retrieved results will lack semantic coherence until the embedding service recovers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →