# Semantica Codebase Submodules: Modular Architecture Guide

> Explore how the Semantica codebase utilizes 18 submodules for a modular architecture. Discover the stages of its knowledge-graph pipeline from data ingestion to export.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: architecture
- Published: 2026-09-08

---

**The Semantica codebase is organized into 18 specialized submodules under the top-level `semantica` package, each encapsulating a distinct stage of the knowledge-graph pipeline from data ingestion through graph construction, reasoning, and export.**

The open-source Semantica framework (`semantica-agi/semantica`) implements a pipeline-based architecture where each functional layer lives in its own submodule. This modular design allows developers to import only the components they need, minimizing dependencies while supporting everything from raw document ingestion to persistent agent memory.

## Semantica Submodule Overview

According to the official [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md) and [`choose-your-module.md`](https://github.com/semantica-agi/semantica/blob/main/choose-your-module.md) documentation, Semantica follows a linear data flow that moves from raw sources through extraction, enrichment, and storage. The architecture diagram illustrates this progression: Sources → Ingest → Parse → Normalize → Split → Extract → Conflict Detection → Deduplication → KG Construction → Enriched KG (ontology + reasoning + provenance + context) → Storage (vector + graph) → Export / Visualization / Services.

Each submodule exposes its public API through an [`__init__.py`](https://github.com/semantica-agi/semantica/blob/main/__init__.py) file (with the exception of `mcp_server`), making the import pattern consistent across the framework.

## Ingestion and Processing Submodules

The initial pipeline stages handle raw data acquisition and preparation across four submodules:

**`semantica.ingest`** loads source material via `FileIngestor`, `WebIngestor`, and `DBIngestor` classes. The entry point is [`semantica/ingest/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/__init__.py).

**`semantica.parse`** converts raw bytes and HTML into structured text objects using `DocumentParser`. The entry point is [`semantica/parse/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/parse/__init__.py).

**`semantica.normalize`** standardizes text casing, dates, numbers, and encodings through `TextNormalizer` and `EntityNormalizer`. The entry point is [`semantica/normalize/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/normalize/__init__.py).

**`semantica.split`** chunks text for downstream processing via the `TextSplitter` class. The entry point is [`semantica/split/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/__init__.py).

## Graph Construction Submodules

The extraction and graph-building submodules transform processed text into structured knowledge:

**`semantica.semantic_extract`** pulls entities, relations, and RDF triplets using `NERExtractor`, `RelationExtractor`, and `TripletExtractor`. The entry point is [`semantica/semantic_extract/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/__init__.py).

**`semantica.conflicts`** detects contradictory facts across sources with `ConflictDetector` and `ConflictResolver`. The entry point is [`semantica/conflicts/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/__init__.py).

**`semantica.deduplication`** collapses duplicate entities and merges metadata via `DuplicateDetector` and `EntityMerger`. The entry point is [`semantica/deduplication/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/deduplication/__init__.py).

**`semantica.kg`** assembles nodes and edges, resolves identities, and adds temporal validity using `GraphBuilder` and `TemporalGraphQuery`. The entry point is [`semantica/kg/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/__init__.py).

## Enrichment and Reasoning Submodules

These submodules add semantic layers and analytical capabilities:

**`semantica.ontology`** generates and validates OWL/SHACL schemas through `OntologyGenerator` and `OntologyValidator`. The entry point is [`semantica/ontology/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ontology/__init__.py).

**`semantica.reasoning`** applies rule-based, Datalog, and SPARQL reasoning via the `Reasoner` and `GraphReasoner` classes. The entry point is [`semantica/reasoning/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/reasoning/__init__.py).

**`semantica.provenance`** tracks source lineage, extraction methods, and checksums using W3C PROV-O standards through `ProvenanceManager`. The entry point is [`semantica/provenance/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/__init__.py).

**`semantica.context`** provides persistent agent memory, decision tracking, and causal-chain analysis via `AgentContext` and `ContextGraph`. The entry point is [`semantica/context/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context/__init__.py).

## Storage Submodules

Semantica abstracts multiple backend systems through dedicated storage submodules:

**`semantica.vector_store`** provides a unified façade for dense embeddings supporting FAISS, Qdrant, Weaviate, Milvus, Pinecone, and PgVector via the `VectorStore` class. The entry point is [`semantica/vector_store/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/__init__.py).

**`semantica.graph_store`** persists knowledge graphs in Neo4j, FalkorDB, and Amazon Neptune using `Neo4jStore` and `FalkorDBStore`. The entry point is [`semantica/graph_store/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/__init__.py).

## Export and Integration Submodules

The final layer handles output formatting and external integration:

**`semantica.export`** serializes graphs to RDF Turtle, JSON-LD, N-Triples, Parquet, and Cypher via `RDFExporter` and `ParquetExporter`. The entry point is [`semantica/export/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/export/__init__.py).

**`semantica.visualization`** renders interactive visualizations for knowledge graphs and embeddings through `KGVisualizer` and `EmbeddingVisualizer`. The entry point is [`semantica/visualization/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/__init__.py).

**`semantica.core`** registers custom components and extends the framework via `PluginRegistry`. The entry point is [`semantica/core/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/core/__init__.py).

**`semantica.mcp_server`** exposes functionality via the multi-client protocol for Claude Desktop, Cursor, and VS Code through `semantica-mcp`. Unlike other submodules, this resides in a single file at [`semantica/mcp_server.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/mcp_server.py).

## Working with Individual Submodules

Each submodule can be imported independently. The following examples demonstrate common workflows using specific submodules.

```python

# 1️⃣ Ingest → Parse → Extract → Build a knowledge graph

from semantica.ingest import FileIngestor               # ← ingest submodule

from semantica.parse import DocumentParser              # ← parse submodule

from semantica.semantic_extract import NERExtractor, RelationExtractor  # ← semantic_extract

from semantica.kg import GraphBuilder                    # ← kg submodule

raw_docs = FileIngestor().ingest("report.pdf")
parsed   = DocumentParser().parse_document("report.pdf")
entities = NERExtractor(method="pattern").extract(parsed)
relations = RelationExtractor(method="rule").extract(parsed, entities=entities)

kg = GraphBuilder(merge_entities=True).build(
    sources=[{"entities": entities, "relationships": relations}]
)
print(f"Graph has {len(kg.nodes)} nodes and {len(kg.edges)} edges")

```

```python

# 2️⃣ Add persistent memory and decision tracking (Context submodule)

from semantica.context import AgentContext, ContextGraph
from semantica.vector_store import VectorStore

ctx = AgentContext(
    vector_store=VectorStore(backend="faiss", dimension=768),
    knowledge_graph=ContextGraph(advanced_analytics=True),
    decision_tracking=True,
)

ctx.store("Apple Inc. was co‑founded by Steve Jobs in 1976.")
decision_id = ctx.record_decision(
    category="model_selection",
    scenario="Select LLM for reasoning pipeline",
    reasoning="GPT‑4 shows higher benchmark scores",
    outcome="selected_gpt4",
    confidence=0.92,
)
print(f"Recorded decision {decision_id}")

```

```python

# 3️⃣ Export the enriched graph to RDF Turtle (Export submodule)

from semantica.export import RDFExporter

exporter = RDFExporter()
exporter.export(kg, "my_graph.ttl", format="turtle")
print("RDF export complete")

```

## Summary

- The Semantica framework organizes functionality into 18 discrete submodules under the `semantica` package namespace.
- The pipeline flows from `ingest` → `parse` → `normalize` → `split` → `semantic_extract` → `conflicts` → `deduplication` → `kg`.
- Enrichment modules (`ontology`, `reasoning`, `provenance`, `context`) layer additional semantics atop the core graph.
- Storage backends are abstracted through `vector_store` and `graph_store` submodules.
- Each submodule exposes its API through [`__init__.py`](https://github.com/semantica-agi/semantica/blob/main/__init__.py) files (except `mcp_server`), enabling selective imports to minimize dependency footprints.

## Frequently Asked Questions

### What is the primary entry point for each Semantica submodule?

Each submodule exposes its public classes and functions through the [`__init__.py`](https://github.com/semantica-agi/semantica/blob/main/__init__.py) file located at `semantica/<submodule>/__init__.py`. For example, [`semantica/ingest/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/__init__.py) provides access to `FileIngestor` and `WebIngestor`, while [`semantica/kg/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/__init__.py) exposes `GraphBuilder` and `TemporalGraphQuery`. The `mcp_server` submodule differs slightly, residing directly at [`semantica/mcp_server.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/mcp_server.py).

### Can I use Semantica submodules independently without loading the entire framework?

Yes. The modular architecture allows selective imports of only the submodules required for your use case. Import `semantica.ingest` without pulling in visualization or reasoning dependencies, keeping your runtime footprint minimal and dependency trees shallow.

### How does the data flow between Semantica submodules?

Data flows linearly from ingestion through export as documented in [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md). Raw sources enter through `semantica.ingest`, pass through parsing and normalization stages, undergo extraction and conflict resolution in `semantica.semantic_extract` and `semantica.conflicts`, then feed into `semantica.kg` for graph construction before persisting in `semantica.vector_store` or `semantica.graph_store`.

### Which submodule handles agent memory and decision tracking?

The `semantica.context` submodule provides persistent agent memory and decision intelligence through the `AgentContext` and `ContextGraph` classes. It tracks causal chains, records decisions with confidence scores, and maintains temporal context for AI agents operating over the knowledge graph.