# Where to Find the Architecture Documentation for Semantica: Complete Pipeline Reference

> Find Semantica architecture documentation in ARCHITECTURE.md. Explore detailed Mermaid diagrams visualizing the complete data pipeline from ingestion to decision intelligence.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: architecture
- Published: 2026-09-08

---

**The complete architecture documentation for Semantica is stored in [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md) at the repository root, containing Mermaid diagrams that visualize the full data pipeline from raw ingestion to knowledge graph construction and the decision intelligence lifecycle.**

The semantica-agi/semantica repository implements a comprehensive semantic data platform with modular Python packages. Understanding the architecture documentation is essential for developers extending the pipeline or integrating specific components like vector stores or graph databases.

## Location of the Architecture Documentation

The authoritative source for system architecture resides in [[`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md)](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md) at the repository root. This file contains two primary Mermaid diagrams: one mapping the complete data processing pipeline across thirteen distinct layers, and another detailing the Decision Intelligence Lifecycle. According to the semantica-agi/semantica source code, these diagrams illustrate how data flows from raw sources through ingestion, parsing, normalization, and semantic extraction into a persistent Knowledge Graph.

## Data Pipeline Architecture Layers

The architecture organizes functionality into thirteen sequential layers, each mapped to specific Python packages under the `semantica/` directory.

### Sources and Ingestion Layer

The **Sources** layer handles raw data ingestion from files, web pages, databases, cloud services, streams, and developer artifacts. The corresponding `semantica.ingest` package implements specialized ingestors including `FileIngestor`, `WebIngestor`, and `DBIngestor`, which transform each source into a uniform "raw document" representation.

### Parsing and Normalization

The **Parse** layer (`semantica.parse`) converts raw documents into structured formats, handling text, code, email, and other document types. The **Normalize** layer (`semantica.normalize`) subsequently cleans and standardizes content, normalizing text entities, dates, and numbers into consistent representations.

### Splitting and Semantic Extraction

The **Split** layer (`semantica.split`) breaks normalized data into granular pieces using entity-aware, graph-based, or ontology-aware strategies. The **Extract** layer (`semantica.semantic_extract`) applies semantic extractors including Named Entity Recognition (NER), relation extraction, event detection, and coreference resolution to identify meaningful patterns.

### Conflict Resolution and Deduplication

Before Knowledge Graph construction, the pipeline runs **Conflict Detection** (`semantica.conflicts`) to identify and resolve contradictory information across sources. The **Deduplication** layer (`semantica.deduplication`) removes duplicate entities and merges redundant records to ensure graph integrity.

### Knowledge Graph Construction

The **KG Construction** layer (`semantica.kg`) builds the core Knowledge Graph structure with nodes, edges, temporal facts, and provenance metadata. This layer outputs a graph representation ready for enrichment and storage.

### Intelligence and Storage Infrastructure

The **Intelligence Layer** (`semantica.ontology`, `semantica.reasoning`, `semantica.provenance`, `semantica.context`) enriches the KG with ontological schemas, reasoning engines, provenance tracking, and contextual decision graphs. The **Storage** layer persists this enriched data through adapters in `semantica.vector_store` (supporting FAISS, Qdrant, PgVector) and `semantica.graph_store` (supporting Neo4j, Amazon Neptune).

### Output Interfaces and Services

The **Outputs** layer (`semantica.export`, `semantica.visualization`) exports data as RDF, JSON-LD, CSV, and visualization formats. The **Services** layer (`semantica.services`) exposes the platform via REST APIs, MCP server protocols, CLI commands, and the Knowledge Explorer UI, with entry points documented in [`README.md`](https://github.com/semantica-agi/semantica/blob/main/README.md).

## Decision Intelligence Lifecycle

Beyond the data pipeline, [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md) diagrams the Decision Intelligence Lifecycle. This process records decisions, links them causally within the knowledge graph, queries similar historical decisions, governs through automated policy checks, and exports audit trails in PROV-O, CSV, or JSON formats. This lifecycle integrates with the Intelligence Layer to provide traceable, auditable decision-making capabilities.

## Implementation Reference: Navigating the Source Code

Each architectural layer maps directly to specific source directories:

- `semantica/ingest/` - Ingestor implementations for files, web, databases, and streams
- `semantica/parse/` - Document parsing logic for various formats
- `semantica/normalize/` - Text and entity normalization utilities
- `semantica/split/` - Chunking and splitting strategies
- `semantica/semantic_extract/` - NER and relation extraction components
- `semantica/kg/` - Knowledge graph building and enrichment
- `semantica/vector_store/` - Vector database adapters (FAISS, Qdrant, PgVector)
- `semantica/graph_store/` - Graph database adapters (Neo4j, Amazon Neptune)
- `semantica/export/` - RDF, CSV, and JSON-LD export utilities
- `semantica/visualization/` - KG and embedding visualization tools

## Practical Example: Running the Full Pipeline

The following Python code exercises the complete pipeline from ingestion to export, corresponding to the architecture layers defined in [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md):

```python
from semantica.ingest import FileIngestor
from semantica.parse import DocumentParser
from semantica.normalize import TextNormalizer
from semantica.split import EntityAwareSplitter
from semantica.semantic_extract import NamedEntityRecognizer
from semantica.kg import GraphBuilder
from semantica.vector_store import PgVectorStore
from semantica.export import RDFExporter

# 1️⃣ Ingest a PDF file

raw_docs = FileIngestor().ingest_path("data/report.pdf")

# 2️⃣ Parse the raw document

parsed = DocumentParser().parse(raw_docs)

# 3️⃣ Normalize the text

normalized = TextNormalizer().normalize(parsed)

# 4️⃣ Split into entity‑aware chunks

chunks = EntityAwareSplitter().split(normalized)

# 5️⃣ Extract named entities

entities = NamedEntityRecognizer().extract(chunks)

# 6️⃣ Build a knowledge graph

kg = GraphBuilder().build(entities)

# 7️⃣ Persist in a PostgreSQL‑backed vector store

vector_store = PgVectorStore(connection_string="postgresql://user:pwd@host/db")
vector_store.upsert(kg)

# 8️⃣ Export as RDF Turtle for downstream consumption

RDFExporter().export(kg, "output.ttl")

```

Running this code executes the full sequence from **ingest** → **parse** → **normalize** → **split** → **extract** → **KG construction** → **storage** → **export**.

## Summary

- The primary architecture documentation for Semantica is located in [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md) at the repository root.
- The documentation contains two Mermaid diagrams: the full data pipeline and the Decision Intelligence Lifecycle.
- The pipeline comprises thirteen layers from Sources through Services, mapped to specific `semantica.*` packages.
- Key storage adapters support FAISS, Qdrant, PgVector for vectors and Neo4j, Amazon Neptune for graphs.
- The Decision Intelligence Lifecycle provides causal decision linking, policy governance, and PROV-O audit trails.

## Frequently Asked Questions

### Where is the architecture documentation located in the Semantica repository?

The architecture documentation is located in the [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md) file at the root of the semantica-agi/semantica repository. This file contains comprehensive Mermaid diagrams illustrating the data pipeline and decision intelligence workflows.

### What does the Semantica architecture documentation contain?

The documentation contains two primary Mermaid diagrams. The first maps the thirteen-layer data pipeline from ingestion to knowledge graph construction. The second describes the Decision Intelligence Lifecycle, covering decision recording, causal linking, policy governance, and audit trail export formats including PROV-O and JSON.

### Which Python packages correspond to the architecture layers?

Each layer maps to a specific package: `semantica.ingest` for data ingestion, `semantica.parse` and `semantica.normalize` for processing, `semantica.semantic_extract` for NER and relation extraction, `semantica.kg` for graph construction, and `semantica.vector_store`/`semantica.graph_store` for persistence. The `semantica.services` package implements the external APIs and interfaces.

### How does the Decision Intelligence Lifecycle work in Semantica?

The Decision Intelligence Lifecycle records decisions within the knowledge graph, links them causally to supporting evidence, enables querying of similar historical decisions, enforces governance through automated policy checks, and exports complete audit trails in PROV-O, CSV, or JSON formats for compliance and verification.