Key Files for Understanding Semantica's Architecture: From Ingestion to Knowledge Graphs

The core architecture of Semantica AGI is organized into ten logical layers—from ingestion through knowledge-graph construction to visualization—with critical entry points located in semantica/pipeline/pipeline_builder.py, semantica/kg/knowledge_graph.py, and ARCHITECTURE.md.

Semantica is an open-source data-to-insight pipeline designed to transform raw content into actionable knowledge graphs through systematic ingestion, extraction, and reasoning. Understanding the key files for understanding Semantica's architecture requires tracing the modular flow from source data to enriched outputs, anchored by the repository's central ARCHITECTURE.md and a hierarchy of specialized Python modules.

ARCHITECTURE.md: The Architectural Blueprint

Before diving into source code, examine ARCHITECTURE.md at the repository root. This document contains a Mermaid diagram that visualizes the complete data flow: Sources → Ingestors → Raw Documents → Parse → Normalize → Split → Extract → Conflict Detection → Deduplication → KG Construction → Enriched KG → Vector / Graph Store → Export / Visualise / Services.

Ingestion Layer: Connecting to Data Sources

The Ingestion layer pulls data from files, web pages, databases, and streams. Three files anchor this stage:

Document Processing Pipeline

After ingestion, documents move through parsing, normalization, splitting, and extraction phases:

Knowledge Graph Construction

The KG layer transforms extracted triples into a queryable graph structure:

Intelligence and Reasoning Layer

Semantica supports multiple reasoning backends for inference over the knowledge graph:

Storage Backends

The storage layer persists enriched data in both vector and graph databases:

Pipeline Orchestration

Two files coordinate the end-to-end execution:

Output and Visualization

Final outputs are generated through specialized export and visualization modules:

Practical Implementation Examples

Building a Complete Ingestion-to-KG Pipeline

The following example demonstrates how PipelineBuilder orchestrates the full data flow from file ingestion to knowledge graph visualization:

from semantica.pipeline.pipeline_builder import PipelineBuilder
from semantica.ingest.file_ingestor import FileIngestor
from semantica.visualization.kg_visualizer import KGVisualizer

# Assemble the pipeline

pipeline = (
    PipelineBuilder()
    .add_ingestor(FileIngestor(source_path="data/example.pdf"))
    .add_parser()           # uses semantica.parse.DocumentParser internally

    .add_normalizer()
    .add_splitter()
    .add_extractor()
    .add_kg_builder()
    .build()
)

# Execute and visualize

kg = pipeline.run()
visualizer = KGVisualizer()
visualizer.render(kg, output_path="kg_viz.html")

Querying Vector Stores for Decision Intelligence

This snippet shows how DecisionRecorder and VectorStore work together to enable similarity-based decision retrieval:

from semantica.vector_store.vector_store import VectorStore
from semantica.context.decision_recorder import DecisionRecorder

# Record a decision

recorder = DecisionRecorder()
decision_id = recorder.record_decision(
    category="risk",
    scenario="data-leak",
    reasoning="High-severity, external exposure",
    outcome="mitigate",
)

# Search for similar past decisions

store = VectorStore()
results = store.search_similar(decision_id, top_k=5)

for r in results:
    print(r.metadata["category"], r.score)

Running Rule-Based Reasoning

Use ReteEngine to apply forward-chaining rules over the knowledge graph:

from semantica.kg.knowledge_graph import KnowledgeGraph
from semantica.reasoning.rete_engine import ReteEngine

kg = KnowledgeGraph.load("path/to/graph.db")
engine = ReteEngine(kg)

# Define a risk-flagging rule

engine.add_rule(
    condition=lambda n: n.get("risk_score", 0) > 7,
    action=lambda n: n.set("flagged", True)
)

engine.run()
print("Flagged nodes:", [n.id for n in kg.nodes if n.get("flagged")])

Summary

Frequently Asked Questions

What is the entry point for building a data pipeline in Semantica?

The semantica/pipeline/pipeline_builder.py module provides the PipelineBuilder class, which offers a fluent API for declaratively assembling pipelines. This builder connects ingestors, parsers, normalizers, extractors, and knowledge graph constructors into a single executable workflow.

Semantica abstracts vector storage through semantica/vector_store/vector_store.py, which defines the interface for all backend implementations. The default local implementation uses semantica/vector_store/faiss_store.py for FAISS-based similarity search, while the system also supports Qdrant, Milvus, and other vector databases through the abstract API.

Which files manage conflict detection and deduplication in the knowledge graph?

The semantica/conflicts/__init__.py module identifies contradictory facts before they enter the knowledge graph, while semantica/deduplication/__init__.py removes duplicate information. Both operate during the transition from extraction to knowledge graph construction, ensuring graph integrity before storage in semantica/kg/knowledge_graph.py.

Where is the reasoning logic implemented in Semantica?

Reasoning is coordinated through semantica/reasoning/reasoner.py, which serves as the high-level entry point. Specific implementations include semantica/reasoning/datalog_reasoner.py for Datalog-style inference and semantica/reasoning/rete_engine.py for forward-chaining rule execution over the knowledge graph.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →