Main Components of the Semantica AI Pipeline: A Complete Technical Breakdown
The Semantica AI pipeline consists of ten core building blocks: a fluent DSL for pipeline construction, a DAG validator, an execution orchestrator, JSON serialization with versioning, a decision-embedding pipeline with graph enrichment, pluggable vector-store backends, reusable MCP tools, visualization utilities, LangChain/Google ADK integrations, and background worker infrastructure.
The semantica-agi/semantica repository provides a flexible, extensible framework for assembling, validating, and executing AI pipelines entirely from Python code. Understanding the main components of the Semantica AI pipeline enables developers to build complex workflows that combine data ingestion, LLM reasoning, vector search, and graph analytics in a single, testable architecture.
Pipeline Construction DSL
At the heart of the system lies a domain-specific language (DSL) for declaratively defining workflows. The entry point is PipelineBuilder, implemented in [semantica/pipeline/pipeline_builder.py](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/pipeline_builder.py), which exposes a fluent API to add steps, define dependencies, set parallelism levels, and register custom handlers.
Each step is represented by a PipelineStep dataclass that captures metadata including name, type, configuration, dependencies, status, and results. The builder aggregates these into a Pipeline dataclass that canonically represents the ordered collection of steps plus pipeline-wide configuration and versioning metadata.
from semantica.pipeline import PipelineBuilder
builder = PipelineBuilder()
builder.register_step_handler("file_ingest", my_file_ingest)
builder.register_step_handler("document_parse", my_document_parser)
pipeline = (
builder
.add_step("ingest", "file_ingest", path="data/documents/")
.add_step("parse", "document_parse")
.connect_steps("ingest", "parse")
.set_parallelism(level=4)
.build(name="doc_embedding_pipeline")
)
Validation & Optimization
Before execution, the PipelineValidator—located in [semantica/pipeline/pipeline_validator.py](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/pipeline_validator.py)—performs static analysis on the directed acyclic graph (DAG). It detects cycles, identifies missing step handlers, validates parallel-safe flags, and flags structural misconfigurations that would cause runtime failures.
Execution Engine
The ExecutionEngine in [semantica/pipeline/execution_engine.py](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/execution_engine.py) orchestrates the actual run. It respects step dependencies, schedules parallel-safe stages concurrently (up to the configured parallelism limit), and tracks progress via an internal progress tracker.
from semantica.pipeline.execution_engine import ExecutionEngine
engine = ExecutionEngine()
result = engine.execute_pipeline(pipeline)
print("Pipeline finished – produced", len(result), "embeddings")
Serialization & Versioning
Pipelines are fully serializable via PipelineSerializer, enabling reproducibility and storage. The implementation in [semantica/pipeline/pipeline_builder.py](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/pipeline_builder.py) supports round-trip conversion to JSON (and other formats) with optional version metadata attachment.
# Serialize to JSON for storage or transmission
json_repr = builder.serialize(format="json")
# Reconstruct later
serializer = PipelineSerializer()
pipeline = serializer.deserialize_pipeline(json_repr)
serializer.version_pipeline(pipeline, {"version": "2.1", "author": "alice"})
Decision-Centric Embedding Pipeline
A specialized high-level component, DecisionEmbeddingPipeline ([semantica/vector_store/decision_embedding_pipeline.py](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/decision_embedding_pipeline.py)), embeds decision objects into vector stores while optionally enriching them with graph-based algorithms. It supports similarity search, path-finding, centrality analysis, and community detection over the embedded decisions.
from semantica.vector_store.decision_embedding_pipeline import DecisionEmbeddingPipeline
from semantica.vector_store.backends import QdrantVectorStore, Neo4jGraphStore
vector_store = QdrantVectorStore(collection="decisions")
graph_store = Neo4jGraphStore(uri="bolt://localhost:7687", auth=("neo4j", "pwd"))
pipeline = DecisionEmbeddingPipeline(
vector_store=vector_store,
graph_store=graph_store,
auto_embed=True,
semantic_weight=0.7,
structural_weight=0.3,
)
embedding = pipeline.process_decision({
"id": "dec-123",
"text": "Approve purchase of 10 k units",
"entities": ["purchase", "10 k units"]
})
Vector-Store Backends
Semantica abstracts vector storage behind a common façade, shipping adapters for FAISS, Qdrant, Pinecone, Weaviate, Milvus, and pgvector. Each backend implements the store interface used interchangeably by the embedding pipeline. Example implementations include [semantica/vector_store/qdrant_store.py](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/qdrant_store.py) and [semantica/vector_store/faiss_index.py](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/faiss_index.py).
Multi-Component Platform (MCP) Tools
The semantica_mcp package provides reusable tools that wire into any pipeline:
retrieval.py– Retrieves relevant context from vector stores ([semantica_mcp/mcp/tools/retrieval.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/retrieval.py))reasoning.py– Executes LLM reasoning on retrieved data ([semantica_mcp/mcp/tools/reasoning.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/reasoning.py))graph.py– Manipulates and queries property-graph stores ([semantica_mcp/mcp/tools/graph.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/graph.py))extraction.py– Extracts structured entities and relations from text ([semantica_mcp/mcp/tools/extraction.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/extraction.py))decisions.py– Persists decision objects and triggers downstream pipelines ([semantica_mcp/mcp/tools/decisions.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/decisions.py))export.py– Exports pipeline artifacts to JSON/YAML and generates visualizations ([semantica_mcp/mcp/tools/export.py](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/export.py))
Visualization Layer
The framework includes a suite of visualizers for temporal, semantic-network, knowledge-graph, and embedding-space views. The core entry point is [semantica/visualization/visualization_provenance.py](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/visualization_provenance.py), which renders pipeline provenance and result graphs for debugging and reporting.
Integration Packages
Semantica integrates with popular AI orchestration frameworks:
- LangChain – Adapters in
integrations/langchain/(e.g., [vectorstore.py](https://github.com/semantica-agi/semantica/blob/main/integrations/langchain/vectorstore.py)) expose pipelines as LangChain agents' tools and retrievers. - Google ADK – Wrappers such as
kg_tools.pyanddecision_tools.pysurface pipeline steps as Google Assistant actions.
Core Runtime & Utilities
Supporting infrastructure includes worker.py ([semantica/worker.py](https://github.com/semantica-agi/semantica/blob/main/semantica/worker.py)), which serves as the entry point for background pipeline execution, and a utils module providing logging, custom exception types, and progress tracking utilities used across all other components.
Summary
- PipelineBuilder DSL – Fluent API in
pipeline_builder.pyfor constructing step-based workflows. - PipelineValidator – Static DAG analysis catching cycles and misconfigurations before runtime.
- ExecutionEngine – Dependency-aware orchestrator with parallel execution support.
- PipelineSerializer – JSON serialization and versioning for reproducible pipelines.
- DecisionEmbeddingPipeline – Specialized pipeline combining vector embedding with graph analytics.
- Vector-Store Backends – Unified adapters for Qdrant, FAISS, Pinecone, and other providers.
- MCP Tools – Reusable retrieval, reasoning, graph, extraction, decision, and export utilities.
- Visualization Layer – Provenance and result graphing via
visualization_provenance.py. - Integration Packages – LangChain and Google ADK compatibility layers.
- Core Runtime – Background workers and shared utilities in
worker.pyandutils.
Frequently Asked Questions
What is the PipelineBuilder in Semantica?
PipelineBuilder is the primary DSL entry point located in semantica/pipeline/pipeline_builder.py. It provides a fluent interface to register custom handlers, add named steps, connect dependencies, configure parallelism levels, and compile everything into an immutable Pipeline dataclass ready for validation and execution.
How does the ExecutionEngine handle parallel execution?
The ExecutionEngine respects the parallelism level set via the builder and executes steps concurrently only when they are marked as parallel-safe and have no unresolved dependencies. It uses an internal progress tracker to monitor stage completion and ensures dependency ordering is strictly maintained according to the DAG defined in the pipeline.
What vector stores does the Semantica AI pipeline support?
Semantica ships with adapters for FAISS, Qdrant, Pinecone, Weaviate, Milvus, and pgvector. Each backend implements a common interface, allowing the DecisionEmbeddingPipeline and retrieval tools to switch storage providers without code changes beyond instantiation.
How do MCP tools integrate with the pipeline?
Multi-Component Platform (MCP) tools are reusable Python modules in semantica_mcp/mcp/tools/ that implement specific operations like vector retrieval or entity extraction. You register them as step handlers via PipelineBuilder.register_step_handler(), enabling any MCP tool to be invoked as a named step within your pipeline DAG.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →