Where to Find the Architecture Documentation for Semantica: Complete Pipeline Reference

The complete architecture documentation for Semantica is stored in ARCHITECTURE.md at the repository root, containing Mermaid diagrams that visualize the full data pipeline from raw ingestion to knowledge graph construction and the decision intelligence lifecycle.

The semantica-agi/semantica repository implements a comprehensive semantic data platform with modular Python packages. Understanding the architecture documentation is essential for developers extending the pipeline or integrating specific components like vector stores or graph databases.

Location of the Architecture Documentation

The authoritative source for system architecture resides in [ARCHITECTURE.md](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md) at the repository root. This file contains two primary Mermaid diagrams: one mapping the complete data processing pipeline across thirteen distinct layers, and another detailing the Decision Intelligence Lifecycle. According to the semantica-agi/semantica source code, these diagrams illustrate how data flows from raw sources through ingestion, parsing, normalization, and semantic extraction into a persistent Knowledge Graph.

Data Pipeline Architecture Layers

The architecture organizes functionality into thirteen sequential layers, each mapped to specific Python packages under the semantica/ directory.

Sources and Ingestion Layer

The Sources layer handles raw data ingestion from files, web pages, databases, cloud services, streams, and developer artifacts. The corresponding semantica.ingest package implements specialized ingestors including FileIngestor, WebIngestor, and DBIngestor, which transform each source into a uniform "raw document" representation.

Parsing and Normalization

The Parse layer (semantica.parse) converts raw documents into structured formats, handling text, code, email, and other document types. The Normalize layer (semantica.normalize) subsequently cleans and standardizes content, normalizing text entities, dates, and numbers into consistent representations.

Splitting and Semantic Extraction

The Split layer (semantica.split) breaks normalized data into granular pieces using entity-aware, graph-based, or ontology-aware strategies. The Extract layer (semantica.semantic_extract) applies semantic extractors including Named Entity Recognition (NER), relation extraction, event detection, and coreference resolution to identify meaningful patterns.

Conflict Resolution and Deduplication

Before Knowledge Graph construction, the pipeline runs Conflict Detection (semantica.conflicts) to identify and resolve contradictory information across sources. The Deduplication layer (semantica.deduplication) removes duplicate entities and merges redundant records to ensure graph integrity.

Knowledge Graph Construction

The KG Construction layer (semantica.kg) builds the core Knowledge Graph structure with nodes, edges, temporal facts, and provenance metadata. This layer outputs a graph representation ready for enrichment and storage.

Intelligence and Storage Infrastructure

The Intelligence Layer (semantica.ontology, semantica.reasoning, semantica.provenance, semantica.context) enriches the KG with ontological schemas, reasoning engines, provenance tracking, and contextual decision graphs. The Storage layer persists this enriched data through adapters in semantica.vector_store (supporting FAISS, Qdrant, PgVector) and semantica.graph_store (supporting Neo4j, Amazon Neptune).

Output Interfaces and Services

The Outputs layer (semantica.export, semantica.visualization) exports data as RDF, JSON-LD, CSV, and visualization formats. The Services layer (semantica.services) exposes the platform via REST APIs, MCP server protocols, CLI commands, and the Knowledge Explorer UI, with entry points documented in README.md.

Decision Intelligence Lifecycle

Beyond the data pipeline, ARCHITECTURE.md diagrams the Decision Intelligence Lifecycle. This process records decisions, links them causally within the knowledge graph, queries similar historical decisions, governs through automated policy checks, and exports audit trails in PROV-O, CSV, or JSON formats. This lifecycle integrates with the Intelligence Layer to provide traceable, auditable decision-making capabilities.

Implementation Reference: Navigating the Source Code

Each architectural layer maps directly to specific source directories:

  • semantica/ingest/ - Ingestor implementations for files, web, databases, and streams
  • semantica/parse/ - Document parsing logic for various formats
  • semantica/normalize/ - Text and entity normalization utilities
  • semantica/split/ - Chunking and splitting strategies
  • semantica/semantic_extract/ - NER and relation extraction components
  • semantica/kg/ - Knowledge graph building and enrichment
  • semantica/vector_store/ - Vector database adapters (FAISS, Qdrant, PgVector)
  • semantica/graph_store/ - Graph database adapters (Neo4j, Amazon Neptune)
  • semantica/export/ - RDF, CSV, and JSON-LD export utilities
  • semantica/visualization/ - KG and embedding visualization tools

Practical Example: Running the Full Pipeline

The following Python code exercises the complete pipeline from ingestion to export, corresponding to the architecture layers defined in ARCHITECTURE.md:

from semantica.ingest import FileIngestor
from semantica.parse import DocumentParser
from semantica.normalize import TextNormalizer
from semantica.split import EntityAwareSplitter
from semantica.semantic_extract import NamedEntityRecognizer
from semantica.kg import GraphBuilder
from semantica.vector_store import PgVectorStore
from semantica.export import RDFExporter

# 1️⃣ Ingest a PDF file

raw_docs = FileIngestor().ingest_path("data/report.pdf")

# 2️⃣ Parse the raw document

parsed = DocumentParser().parse(raw_docs)

# 3️⃣ Normalize the text

normalized = TextNormalizer().normalize(parsed)

# 4️⃣ Split into entity‑aware chunks

chunks = EntityAwareSplitter().split(normalized)

# 5️⃣ Extract named entities

entities = NamedEntityRecognizer().extract(chunks)

# 6️⃣ Build a knowledge graph

kg = GraphBuilder().build(entities)

# 7️⃣ Persist in a PostgreSQL‑backed vector store

vector_store = PgVectorStore(connection_string="postgresql://user:pwd@host/db")
vector_store.upsert(kg)

# 8️⃣ Export as RDF Turtle for downstream consumption

RDFExporter().export(kg, "output.ttl")

Running this code executes the full sequence from ingest → parse → normalize → split → extract → KG construction → storage → export.

Summary

  • The primary architecture documentation for Semantica is located in ARCHITECTURE.md at the repository root.
  • The documentation contains two Mermaid diagrams: the full data pipeline and the Decision Intelligence Lifecycle.
  • The pipeline comprises thirteen layers from Sources through Services, mapped to specific semantica.* packages.
  • Key storage adapters support FAISS, Qdrant, PgVector for vectors and Neo4j, Amazon Neptune for graphs.
  • The Decision Intelligence Lifecycle provides causal decision linking, policy governance, and PROV-O audit trails.

Frequently Asked Questions

Where is the architecture documentation located in the Semantica repository?

The architecture documentation is located in the ARCHITECTURE.md file at the root of the semantica-agi/semantica repository. This file contains comprehensive Mermaid diagrams illustrating the data pipeline and decision intelligence workflows.

What does the Semantica architecture documentation contain?

The documentation contains two primary Mermaid diagrams. The first maps the thirteen-layer data pipeline from ingestion to knowledge graph construction. The second describes the Decision Intelligence Lifecycle, covering decision recording, causal linking, policy governance, and audit trail export formats including PROV-O and JSON.

Which Python packages correspond to the architecture layers?

Each layer maps to a specific package: semantica.ingest for data ingestion, semantica.parse and semantica.normalize for processing, semantica.semantic_extract for NER and relation extraction, semantica.kg for graph construction, and semantica.vector_store/semantica.graph_store for persistence. The semantica.services package implements the external APIs and interfaces.

How does the Decision Intelligence Lifecycle work in Semantica?

The Decision Intelligence Lifecycle records decisions within the knowledge graph, links them causally to supporting evidence, enables querying of similar historical decisions, enforces governance through automated policy checks, and exports complete audit trails in PROV-O, CSV, or JSON formats for compliance and verification.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →