Semantica Codebase Submodules: Modular Architecture Guide

The Semantica codebase is organized into 18 specialized submodules under the top-level semantica package, each encapsulating a distinct stage of the knowledge-graph pipeline from data ingestion through graph construction, reasoning, and export.

The open-source Semantica framework (semantica-agi/semantica) implements a pipeline-based architecture where each functional layer lives in its own submodule. This modular design allows developers to import only the components they need, minimizing dependencies while supporting everything from raw document ingestion to persistent agent memory.

Semantica Submodule Overview

According to the official ARCHITECTURE.md and choose-your-module.md documentation, Semantica follows a linear data flow that moves from raw sources through extraction, enrichment, and storage. The architecture diagram illustrates this progression: Sources → Ingest → Parse → Normalize → Split → Extract → Conflict Detection → Deduplication → KG Construction → Enriched KG (ontology + reasoning + provenance + context) → Storage (vector + graph) → Export / Visualization / Services.

Each submodule exposes its public API through an __init__.py file (with the exception of mcp_server), making the import pattern consistent across the framework.

Ingestion and Processing Submodules

The initial pipeline stages handle raw data acquisition and preparation across four submodules:

semantica.ingest loads source material via FileIngestor, WebIngestor, and DBIngestor classes. The entry point is semantica/ingest/__init__.py.

semantica.parse converts raw bytes and HTML into structured text objects using DocumentParser. The entry point is semantica/parse/__init__.py.

semantica.normalize standardizes text casing, dates, numbers, and encodings through TextNormalizer and EntityNormalizer. The entry point is semantica/normalize/__init__.py.

semantica.split chunks text for downstream processing via the TextSplitter class. The entry point is semantica/split/__init__.py.

Graph Construction Submodules

The extraction and graph-building submodules transform processed text into structured knowledge:

semantica.semantic_extract pulls entities, relations, and RDF triplets using NERExtractor, RelationExtractor, and TripletExtractor. The entry point is semantica/semantic_extract/__init__.py.

semantica.conflicts detects contradictory facts across sources with ConflictDetector and ConflictResolver. The entry point is semantica/conflicts/__init__.py.

semantica.deduplication collapses duplicate entities and merges metadata via DuplicateDetector and EntityMerger. The entry point is semantica/deduplication/__init__.py.

semantica.kg assembles nodes and edges, resolves identities, and adds temporal validity using GraphBuilder and TemporalGraphQuery. The entry point is semantica/kg/__init__.py.

Enrichment and Reasoning Submodules

These submodules add semantic layers and analytical capabilities:

semantica.ontology generates and validates OWL/SHACL schemas through OntologyGenerator and OntologyValidator. The entry point is semantica/ontology/__init__.py.

semantica.reasoning applies rule-based, Datalog, and SPARQL reasoning via the Reasoner and GraphReasoner classes. The entry point is semantica/reasoning/__init__.py.

semantica.provenance tracks source lineage, extraction methods, and checksums using W3C PROV-O standards through ProvenanceManager. The entry point is semantica/provenance/__init__.py.

semantica.context provides persistent agent memory, decision tracking, and causal-chain analysis via AgentContext and ContextGraph. The entry point is semantica/context/__init__.py.

Storage Submodules

Semantica abstracts multiple backend systems through dedicated storage submodules:

semantica.vector_store provides a unified façade for dense embeddings supporting FAISS, Qdrant, Weaviate, Milvus, Pinecone, and PgVector via the VectorStore class. The entry point is semantica/vector_store/__init__.py.

semantica.graph_store persists knowledge graphs in Neo4j, FalkorDB, and Amazon Neptune using Neo4jStore and FalkorDBStore. The entry point is semantica/graph_store/__init__.py.

Export and Integration Submodules

The final layer handles output formatting and external integration:

semantica.export serializes graphs to RDF Turtle, JSON-LD, N-Triples, Parquet, and Cypher via RDFExporter and ParquetExporter. The entry point is semantica/export/__init__.py.

semantica.visualization renders interactive visualizations for knowledge graphs and embeddings through KGVisualizer and EmbeddingVisualizer. The entry point is semantica/visualization/__init__.py.

semantica.core registers custom components and extends the framework via PluginRegistry. The entry point is semantica/core/__init__.py.

semantica.mcp_server exposes functionality via the multi-client protocol for Claude Desktop, Cursor, and VS Code through semantica-mcp. Unlike other submodules, this resides in a single file at semantica/mcp_server.py.

Working with Individual Submodules

Each submodule can be imported independently. The following examples demonstrate common workflows using specific submodules.


# 1️⃣ Ingest → Parse → Extract → Build a knowledge graph

from semantica.ingest import FileIngestor               # ← ingest submodule

from semantica.parse import DocumentParser              # ← parse submodule

from semantica.semantic_extract import NERExtractor, RelationExtractor  # ← semantic_extract

from semantica.kg import GraphBuilder                    # ← kg submodule

raw_docs = FileIngestor().ingest("report.pdf")
parsed   = DocumentParser().parse_document("report.pdf")
entities = NERExtractor(method="pattern").extract(parsed)
relations = RelationExtractor(method="rule").extract(parsed, entities=entities)

kg = GraphBuilder(merge_entities=True).build(
    sources=[{"entities": entities, "relationships": relations}]
)
print(f"Graph has {len(kg.nodes)} nodes and {len(kg.edges)} edges")

# 2️⃣ Add persistent memory and decision tracking (Context submodule)

from semantica.context import AgentContext, ContextGraph
from semantica.vector_store import VectorStore

ctx = AgentContext(
    vector_store=VectorStore(backend="faiss", dimension=768),
    knowledge_graph=ContextGraph(advanced_analytics=True),
    decision_tracking=True,
)

ctx.store("Apple Inc. was co‑founded by Steve Jobs in 1976.")
decision_id = ctx.record_decision(
    category="model_selection",
    scenario="Select LLM for reasoning pipeline",
    reasoning="GPT‑4 shows higher benchmark scores",
    outcome="selected_gpt4",
    confidence=0.92,
)
print(f"Recorded decision {decision_id}")

# 3️⃣ Export the enriched graph to RDF Turtle (Export submodule)

from semantica.export import RDFExporter

exporter = RDFExporter()
exporter.export(kg, "my_graph.ttl", format="turtle")
print("RDF export complete")

Summary

  • The Semantica framework organizes functionality into 18 discrete submodules under the semantica package namespace.
  • The pipeline flows from ingest → parse → normalize → split → semantic_extract → conflicts → deduplication → kg.
  • Enrichment modules (ontology, reasoning, provenance, context) layer additional semantics atop the core graph.
  • Storage backends are abstracted through vector_store and graph_store submodules.
  • Each submodule exposes its API through __init__.py files (except mcp_server), enabling selective imports to minimize dependency footprints.

Frequently Asked Questions

What is the primary entry point for each Semantica submodule?

Each submodule exposes its public classes and functions through the __init__.py file located at semantica/<submodule>/__init__.py. For example, semantica/ingest/__init__.py provides access to FileIngestor and WebIngestor, while semantica/kg/__init__.py exposes GraphBuilder and TemporalGraphQuery. The mcp_server submodule differs slightly, residing directly at semantica/mcp_server.py.

Can I use Semantica submodules independently without loading the entire framework?

Yes. The modular architecture allows selective imports of only the submodules required for your use case. Import semantica.ingest without pulling in visualization or reasoning dependencies, keeping your runtime footprint minimal and dependency trees shallow.

How does the data flow between Semantica submodules?

Data flows linearly from ingestion through export as documented in ARCHITECTURE.md. Raw sources enter through semantica.ingest, pass through parsing and normalization stages, undergo extraction and conflict resolution in semantica.semantic_extract and semantica.conflicts, then feed into semantica.kg for graph construction before persisting in semantica.vector_store or semantica.graph_store.

Which submodule handles agent memory and decision tracking?

The semantica.context submodule provides persistent agent memory and decision intelligence through the AgentContext and ContextGraph classes. It tracks causal chains, records decisions with confidence scores, and maintains temporal context for AI agents operating over the knowledge graph.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →