# Key Files for Understanding Semantica's Architecture: From Ingestion to Knowledge Graphs

> Explore Semantica's architecture by understanding its key files. Discover entry points in pipeline_builder.py and knowledge_graph.py for ingestion to knowledge graphs.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: architecture
- Published: 2026-09-09

---

**The core architecture of Semantica AGI is organized into ten logical layers—from ingestion through knowledge-graph construction to visualization—with critical entry points located in [`semantica/pipeline/pipeline_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/pipeline_builder.py), [`semantica/kg/knowledge_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/knowledge_graph.py), and [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md).**

Semantica is an open-source data-to-insight pipeline designed to transform raw content into actionable knowledge graphs through systematic ingestion, extraction, and reasoning. Understanding the key files for understanding Semantica's architecture requires tracing the modular flow from source data to enriched outputs, anchored by the repository's central [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md) and a hierarchy of specialized Python modules.

## ARCHITECTURE.md: The Architectural Blueprint

Before diving into source code, examine **[`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md)** at the repository root. This document contains a Mermaid diagram that visualizes the complete data flow: **Sources** → **Ingestors** → **Raw Documents** → **Parse** → **Normalize** → **Split** → **Extract** → **Conflict Detection** → **Deduplication** → **KG Construction** → **Enriched KG** → **Vector / Graph Store** → **Export / Visualise / Services**.

## Ingestion Layer: Connecting to Data Sources

The **Ingestion** layer pulls data from files, web pages, databases, and streams. Three files anchor this stage:

- **[`semantica/ingest/registry.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/registry.py)** – Central registry that maintains all available ingestors and routes data to the appropriate handler.
- **[`semantica/ingest/file_ingestor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/file_ingestor.py)** – Processes PDFs, DOCX, TXT, CSV, JSON, Excel, and XML files.
- **[`semantica/ingest/web_ingestor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/web_ingestor.py)** – Crawls web pages, RSS/Atom feeds, and public REST APIs.

## Document Processing Pipeline

After ingestion, documents move through parsing, normalization, splitting, and extraction phases:

- **[`semantica/parse/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/parse/__init__.py)** – Converts raw bytes into structured text.
- **[`semantica/normalize/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/normalize/__init__.py)** – Standardizes dates, numbers, and entities.
- **[`semantica/split/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/__init__.py)** – Decomposes documents into logical chunks for downstream processing.
- **[`semantica/semantic_extract/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/__init__.py)** – Detects named entities, relations, events, and coreference links.

## Knowledge Graph Construction

The KG layer transforms extracted triples into a queryable graph structure:

- **[`semantica/kg/knowledge_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/knowledge_graph.py)** – Core data model managing nodes, edges, temporal facts, and provenance tracking.
- **[`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py)** – Constructs the graph from extracted semantic triples and resolves entity references.
- **[`semantica/conflicts/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/__init__.py)** – Identifies contradictory facts before graph insertion.
- **[`semantica/deduplication/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/deduplication/__init__.py)** – Removes duplicate information to maintain graph integrity.

## Intelligence and Reasoning Layer

Semantica supports multiple reasoning backends for inference over the knowledge graph:

- **[`semantica/reasoning/reasoner.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/reasoning/reasoner.py)** – High-level entry point abstracting all reasoning implementations.
- **[`semantica/reasoning/datalog_reasoner.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/reasoning/datalog_reasoner.py)** – Datalog-style rule engine for logical inference.
- **[`semantica/reasoning/rete_engine.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/reasoning/rete_engine.py)** – Rete-based forward-chaining rule engine for complex pattern matching.
- **[`semantica/ontology/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ontology/__init__.py)** – Handles ontology generation and validation.
- **[`semantica/provenance/manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/manager.py)** – Tracks data lineage and decision provenance.

## Storage Backends

The storage layer persists enriched data in both vector and graph databases:

- **[`semantica/vector_store/vector_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/vector_store.py)** – Abstract API defining interfaces for FAISS, Qdrant, Milvus, and other vector stores.
- **[`semantica/vector_store/faiss_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/faiss_store.py)** – Default local implementation using FAISS for similarity search.
- **[`semantica/graph_store/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/__init__.py)** – Interface for graph databases including Neo4j and Amazon Neptune.

## Pipeline Orchestration

Two files coordinate the end-to-end execution:

- **[`semantica/pipeline/pipeline_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/pipeline_builder.py)** – Declarative builder that stitches ingestion, processing, and storage stages into executable pipelines.
- **[`semantica/pipeline/execution_engine.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/execution_engine.py)** – Executes constructed pipelines with parallelism and failure handling.
- **[`semantica/pipeline/parallelism_manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/parallelism_manager.py)** – Manages concurrent processing across pipeline stages.

## Output and Visualization

Final outputs are generated through specialized export and visualization modules:

- **[`semantica/export/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/export/__init__.py)** – Exports KG data to RDF/JSON-LD, CSV, and Parquet formats.
- **[`semantica/visualization/visualization_provenance.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/visualization_provenance.py)** – Registry for visualization utilities.
- **[`semantica/visualization/kg_visualizer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/kg_visualizer.py)** – Generates interactive HTML/JS visualizations of knowledge graphs.
- **[`semantica/context/decision_recorder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context/decision_recorder.py)** – Records decisions and stores them in the vector store for retrieval.

## Practical Implementation Examples

### Building a Complete Ingestion-to-KG Pipeline

The following example demonstrates how `PipelineBuilder` orchestrates the full data flow from file ingestion to knowledge graph visualization:

```python
from semantica.pipeline.pipeline_builder import PipelineBuilder
from semantica.ingest.file_ingestor import FileIngestor
from semantica.visualization.kg_visualizer import KGVisualizer

# Assemble the pipeline

pipeline = (
    PipelineBuilder()
    .add_ingestor(FileIngestor(source_path="data/example.pdf"))
    .add_parser()           # uses semantica.parse.DocumentParser internally

    .add_normalizer()
    .add_splitter()
    .add_extractor()
    .add_kg_builder()
    .build()
)

# Execute and visualize

kg = pipeline.run()
visualizer = KGVisualizer()
visualizer.render(kg, output_path="kg_viz.html")

```

### Querying Vector Stores for Decision Intelligence

This snippet shows how `DecisionRecorder` and `VectorStore` work together to enable similarity-based decision retrieval:

```python
from semantica.vector_store.vector_store import VectorStore
from semantica.context.decision_recorder import DecisionRecorder

# Record a decision

recorder = DecisionRecorder()
decision_id = recorder.record_decision(
    category="risk",
    scenario="data-leak",
    reasoning="High-severity, external exposure",
    outcome="mitigate",
)

# Search for similar past decisions

store = VectorStore()
results = store.search_similar(decision_id, top_k=5)

for r in results:
    print(r.metadata["category"], r.score)

```

### Running Rule-Based Reasoning

Use `ReteEngine` to apply forward-chaining rules over the knowledge graph:

```python
from semantica.kg.knowledge_graph import KnowledgeGraph
from semantica.reasoning.rete_engine import ReteEngine

kg = KnowledgeGraph.load("path/to/graph.db")
engine = ReteEngine(kg)

# Define a risk-flagging rule

engine.add_rule(
    condition=lambda n: n.get("risk_score", 0) > 7,
    action=lambda n: n.set("flagged", True)
)

engine.run()
print("Flagged nodes:", [n.id for n in kg.nodes if n.get("flagged")])

```

## Summary

- **[`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md)** provides the high-level Mermaid diagram and narrative explaining how data flows through all ten architectural layers.
- **[`semantica/pipeline/pipeline_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/pipeline_builder.py)** serves as the primary entry point for constructing data pipelines that connect ingestion to storage.
- The **Ingestion** layer is centralized in [`semantica/ingest/registry.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/registry.py), with specific implementations in [`file_ingestor.py`](https://github.com/semantica-agi/semantica/blob/main/file_ingestor.py) and [`web_ingestor.py`](https://github.com/semantica-agi/semantica/blob/main/web_ingestor.py).
- **Knowledge graph** construction relies on [`semantica/kg/knowledge_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/knowledge_graph.py) for the data model and [`graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/graph_builder.py) for assembly.
- **Reasoning** capabilities are accessed through [`semantica/reasoning/reasoner.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/reasoning/reasoner.py), with specific implementations including [`datalog_reasoner.py`](https://github.com/semantica-agi/semantica/blob/main/datalog_reasoner.py) and [`rete_engine.py`](https://github.com/semantica-agi/semantica/blob/main/rete_engine.py).
- **Vector storage** is abstracted in [`semantica/vector_store/vector_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/vector_store.py), with FAISS support via [`faiss_store.py`](https://github.com/semantica-agi/semantica/blob/main/faiss_store.py).

## Frequently Asked Questions

### What is the entry point for building a data pipeline in Semantica?

The **[`semantica/pipeline/pipeline_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/pipeline_builder.py)** module provides the `PipelineBuilder` class, which offers a fluent API for declaratively assembling pipelines. This builder connects ingestors, parsers, normalizers, extractors, and knowledge graph constructors into a single executable workflow.

### How does Semantica handle vector storage and similarity search?

Semantica abstracts vector storage through **[`semantica/vector_store/vector_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/vector_store.py)**, which defines the interface for all backend implementations. The default local implementation uses **[`semantica/vector_store/faiss_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/faiss_store.py)** for FAISS-based similarity search, while the system also supports Qdrant, Milvus, and other vector databases through the abstract API.

### Which files manage conflict detection and deduplication in the knowledge graph?

The **[`semantica/conflicts/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/__init__.py)** module identifies contradictory facts before they enter the knowledge graph, while **[`semantica/deduplication/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/deduplication/__init__.py)** removes duplicate information. Both operate during the transition from extraction to knowledge graph construction, ensuring graph integrity before storage in [`semantica/kg/knowledge_graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/knowledge_graph.py).

### Where is the reasoning logic implemented in Semantica?

Reasoning is coordinated through **[`semantica/reasoning/reasoner.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/reasoning/reasoner.py)**, which serves as the high-level entry point. Specific implementations include **[`semantica/reasoning/datalog_reasoner.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/reasoning/datalog_reasoner.py)** for Datalog-style inference and **[`semantica/reasoning/rete_engine.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/reasoning/rete_engine.py)** for forward-chaining rule execution over the knowledge graph.