# Semantica Data Processing Pipeline: 8 Distinct Stages Explained

> Explore Semantica's data processing pipeline: 8 stages including validation, embedding generation, vector combination, and similarity search to create hybrid vector representations.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: deep-dive
- Published: 2026-09-08

---

**Semantica's decision-embedding pipeline processes raw decision records through eight distinct stages—validation, semantic embedding generation, structural embedding generation, vector combination, metadata enrichment, storage, batch processing, and similarity search—to produce enriched hybrid vector representations.**

The `semantica-agi/semantica` repository implements a hybrid architecture that combines language-based semantic analysis with graph-based structural context. Understanding the distinct stages in Semantica's data processing pipeline enables developers to optimize ingestion workflows and leverage advanced knowledge-graph features for decision tracking.

## Stage 1: Validation and Normalization

The pipeline begins by enforcing data integrity through the `_validate_decision_data` method in [`semantica/vector_store/decision_embedding_pipeline.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/decision_embedding_pipeline.py). This stage checks for required fields such as `scenario` and applies sensible defaults for optional fields including `outcome`, `reasoning`, and `confidence`.

Input records that lack mandatory fields are rejected at this boundary, ensuring downstream embedding generation receives consistent, well-formed decision objects.

## Stage 2: Semantic Embedding Generation

Once validated, decision text is converted into dense vectors via `_generate_semantic_embedding`. The implementation concatenates relevant text fields—**scenario**, **reasoning**, **outcome**, and **category**—into a single document representation.

The configured `vector_store.embed` method generates the semantic vector. If the primary embedding service fails, the pipeline automatically falls back to a random vector, ensuring system continuity at the cost of semantic accuracy.

## Stage 3: Structural Embedding Generation (Optional)

When a **graph store** is configured and `use_graph_features=True`, the pipeline invokes `_generate_structural_embedding` to capture relational context. This stage extracts entities from the decision and utilizes `NodeEmbedder.compute_embeddings` (implemented in [`semantica/kg/node_embeddings.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/node_embeddings.py)) to generate Node2Vec-based representations.

Advanced knowledge-graph algorithms enhance these embeddings:

- **Path-finder** (from [`semantica/kg/path_finder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/path_finder.py)) calculates shortest-path contexts
- **Community detector** (from [`semantica/kg/community_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/community_detector.py)) identifies cluster memberships
- **Centrality and connectivity** metrics quantify node importance within the graph

## Stage 4: Combination of Semantic and Structural Vectors

The `_create_combined_embedding` method merges dual-vector representations into a single hybrid embedding. This stage first resizes vectors to identical dimensionalities if they differ, then applies configurable weights (`semantic_weight` and `structural_weight`) to balance linguistic meaning against graph topology.

The default weighting favors semantic content, but users can adjust the ratio to prioritize relational structure for specific use cases.

## Stage 5: Metadata Enrichment

Before persistence, `_enrich_metadata` attaches provenance information to each decision record. Captured metadata includes:

- Pipeline version identifier
- Generation timestamps
- Active weight settings (`semantic_weight`, `structural_weight`)
- Boolean flags indicating structural embedding inclusion

This metadata enables audit trails and facilitates A/B testing of different pipeline configurations.

## Stage 6: Storage

The `_store_embeddings` method persists finalized vectors to the configured vector store. Both the combined hybrid vector and the raw semantic vector are stored alongside enriched metadata, enabling retrieval strategies that query either representation independently.

Storage atomicity ensures that partial failures do not leave orphaned metadata or unreferenced vector entries.

## Stage 7: Batch Processing (Optional)

For high-volume ingestion, `_process_decision_batch` optimizes throughput by pre-generating structural embeddings once per unique entity, then reusing them across the decision batch. This avoids redundant graph computations when processing thousands of related decisions.

Progress tracking is handled by `ProgressTracker` (from [`semantica/utils/progress_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/utils/progress_tracker.py)), which reports completion percentages and estimated time remaining for long-running batch jobs.

## Stage 8: Similarity Search (Optional)

The pipeline supports **hybrid retrieval** via `_find_similar_decisions`. Query decisions traverse the same eight stages to generate comparable vector representations. The system then retrieves candidate embeddings from the vector store and computes hybrid similarity scores combining both semantic proximity and graph structural alignment.

This enables precedent retrieval that respects not just textual similarity but also organizational or causal relationships encoded in the knowledge graph.

## Implementation Example

The following code demonstrates the complete pipeline workflow:

```python

# Initialise the pipeline (semantic vector store + optional graph store)

pipeline = DecisionEmbeddingPipeline(
    vector_store=my_vector_store,
    graph_store=my_graph_store,          # omit to disable structural embeddings

    use_graph_features=True,             # enable KG‑enhanced embeddings

    semantic_weight=0.7,
    structural_weight=0.3,
)

# 1️⃣ Process a single decision

result = pipeline.process_decision({
    "scenario": "Increase credit limit",
    "reasoning": "Customer has good payment history",
    "outcome": "approved",
    "entities": ["Customer", "CreditAccount"],
    "category": "Finance",
})
print(result["combined_embedding"].shape)   # → (384,)

# 2️⃣ Process a batch of decisions

batch = [
    {"scenario": "Add new product line", "entities": ["Product"], "category": "Strategy"},
    {"scenario": "Close under‑performing store", "entities": ["Store"], "category": "Operations"},
]
batch_results = pipeline.process_decision_batch(batch, batch_size=2)

# 3️⃣ Find similar past decisions (hybrid search)

similar = pipeline.find_similar_decisions(
    query_decision={"scenario": "Raise credit limit", "entities": ["Customer"]},
    limit=5,
)
for hit in similar:
    print(hit["metadata"]["scenario"], hit["similarity"])

```

## Summary

- **Validation and Normalization** ensures data integrity before processing via `_validate_decision_data`.
- **Semantic Embedding Generation** converts text fields into dense vectors using `vector_store.embed`.
- **Structural Embedding Generation** optionally extracts graph features via `NodeEmbedder.compute_embeddings`.
- **Vector Combination** merges representations using configurable weights in `_create_combined_embedding`.
- **Metadata Enrichment** adds timestamps and configuration details through `_enrich_metadata`.
- **Storage** persists hybrid vectors via `_store_embeddings`.
- **Batch Processing** optimizes throughput for large datasets using `_process_decision_batch`.
- **Similarity Search** enables hybrid retrieval matching both text and graph structure via `_find_similar_decisions`.

## Frequently Asked Questions

### What is the difference between semantic and structural embeddings in Semantica?

**Semantic embeddings** capture linguistic meaning by encoding concatenated text fields (scenario, reasoning, outcome) through standard vector store models. **Structural embeddings** encode relational context extracted from knowledge graphs using Node2Vec algorithms and graph metrics like centrality and community detection. The pipeline combines both to create hybrid representations that understand both *what* was decided and *how* it relates to other decisions.

### How does Semantica handle invalid decision data during processing?

The `_validate_decision_data` method in [`semantica/vector_store/decision_embedding_pipeline.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/decision_embedding_pipeline.py) acts as a strict gatekeeper. Records missing required fields such as `scenario` are rejected immediately, while optional fields receive default values. This prevents malformed data from corrupting downstream embedding generation or storage operations.

### Can I use Semantica's pipeline without a graph store?

Yes. The graph store is optional. If `graph_store` is omitted or `use_graph_features=False`, the pipeline skips structural embedding generation (Stage 3) and relies solely on semantic vectors. In this mode, `_create_combined_embedding` simply passes through the semantic representation, and similarity search operates on text-based vectors only.

### What happens if the vector store fails to generate an embedding?

The `_generate_semantic_embedding` method includes a fault-tolerance mechanism. If `vector_store.embed` raises an exception or returns an error, the pipeline generates a random fallback vector of the expected dimensionality. This ensures system availability, though retrieved results will lack semantic coherence until the embedding service recovers.