Key Files for Understanding Semantica's Architecture: From Ingestion to Knowledge Graphs
The core architecture of Semantica AGI is organized into ten logical layers—from ingestion through knowledge-graph construction to visualization—with critical entry points located in semantica/pipeline/pipeline_builder.py, semantica/kg/knowledge_graph.py, and ARCHITECTURE.md.
Semantica is an open-source data-to-insight pipeline designed to transform raw content into actionable knowledge graphs through systematic ingestion, extraction, and reasoning. Understanding the key files for understanding Semantica's architecture requires tracing the modular flow from source data to enriched outputs, anchored by the repository's central ARCHITECTURE.md and a hierarchy of specialized Python modules.
ARCHITECTURE.md: The Architectural Blueprint
Before diving into source code, examine ARCHITECTURE.md at the repository root. This document contains a Mermaid diagram that visualizes the complete data flow: Sources → Ingestors → Raw Documents → Parse → Normalize → Split → Extract → Conflict Detection → Deduplication → KG Construction → Enriched KG → Vector / Graph Store → Export / Visualise / Services.
Ingestion Layer: Connecting to Data Sources
The Ingestion layer pulls data from files, web pages, databases, and streams. Three files anchor this stage:
semantica/ingest/registry.py– Central registry that maintains all available ingestors and routes data to the appropriate handler.semantica/ingest/file_ingestor.py– Processes PDFs, DOCX, TXT, CSV, JSON, Excel, and XML files.semantica/ingest/web_ingestor.py– Crawls web pages, RSS/Atom feeds, and public REST APIs.
Document Processing Pipeline
After ingestion, documents move through parsing, normalization, splitting, and extraction phases:
semantica/parse/__init__.py– Converts raw bytes into structured text.semantica/normalize/__init__.py– Standardizes dates, numbers, and entities.semantica/split/__init__.py– Decomposes documents into logical chunks for downstream processing.semantica/semantic_extract/__init__.py– Detects named entities, relations, events, and coreference links.
Knowledge Graph Construction
The KG layer transforms extracted triples into a queryable graph structure:
semantica/kg/knowledge_graph.py– Core data model managing nodes, edges, temporal facts, and provenance tracking.semantica/kg/graph_builder.py– Constructs the graph from extracted semantic triples and resolves entity references.semantica/conflicts/__init__.py– Identifies contradictory facts before graph insertion.semantica/deduplication/__init__.py– Removes duplicate information to maintain graph integrity.
Intelligence and Reasoning Layer
Semantica supports multiple reasoning backends for inference over the knowledge graph:
semantica/reasoning/reasoner.py– High-level entry point abstracting all reasoning implementations.semantica/reasoning/datalog_reasoner.py– Datalog-style rule engine for logical inference.semantica/reasoning/rete_engine.py– Rete-based forward-chaining rule engine for complex pattern matching.semantica/ontology/__init__.py– Handles ontology generation and validation.semantica/provenance/manager.py– Tracks data lineage and decision provenance.
Storage Backends
The storage layer persists enriched data in both vector and graph databases:
semantica/vector_store/vector_store.py– Abstract API defining interfaces for FAISS, Qdrant, Milvus, and other vector stores.semantica/vector_store/faiss_store.py– Default local implementation using FAISS for similarity search.semantica/graph_store/__init__.py– Interface for graph databases including Neo4j and Amazon Neptune.
Pipeline Orchestration
Two files coordinate the end-to-end execution:
semantica/pipeline/pipeline_builder.py– Declarative builder that stitches ingestion, processing, and storage stages into executable pipelines.semantica/pipeline/execution_engine.py– Executes constructed pipelines with parallelism and failure handling.semantica/pipeline/parallelism_manager.py– Manages concurrent processing across pipeline stages.
Output and Visualization
Final outputs are generated through specialized export and visualization modules:
semantica/export/__init__.py– Exports KG data to RDF/JSON-LD, CSV, and Parquet formats.semantica/visualization/visualization_provenance.py– Registry for visualization utilities.semantica/visualization/kg_visualizer.py– Generates interactive HTML/JS visualizations of knowledge graphs.semantica/context/decision_recorder.py– Records decisions and stores them in the vector store for retrieval.
Practical Implementation Examples
Building a Complete Ingestion-to-KG Pipeline
The following example demonstrates how PipelineBuilder orchestrates the full data flow from file ingestion to knowledge graph visualization:
from semantica.pipeline.pipeline_builder import PipelineBuilder
from semantica.ingest.file_ingestor import FileIngestor
from semantica.visualization.kg_visualizer import KGVisualizer
# Assemble the pipeline
pipeline = (
PipelineBuilder()
.add_ingestor(FileIngestor(source_path="data/example.pdf"))
.add_parser() # uses semantica.parse.DocumentParser internally
.add_normalizer()
.add_splitter()
.add_extractor()
.add_kg_builder()
.build()
)
# Execute and visualize
kg = pipeline.run()
visualizer = KGVisualizer()
visualizer.render(kg, output_path="kg_viz.html")
Querying Vector Stores for Decision Intelligence
This snippet shows how DecisionRecorder and VectorStore work together to enable similarity-based decision retrieval:
from semantica.vector_store.vector_store import VectorStore
from semantica.context.decision_recorder import DecisionRecorder
# Record a decision
recorder = DecisionRecorder()
decision_id = recorder.record_decision(
category="risk",
scenario="data-leak",
reasoning="High-severity, external exposure",
outcome="mitigate",
)
# Search for similar past decisions
store = VectorStore()
results = store.search_similar(decision_id, top_k=5)
for r in results:
print(r.metadata["category"], r.score)
Running Rule-Based Reasoning
Use ReteEngine to apply forward-chaining rules over the knowledge graph:
from semantica.kg.knowledge_graph import KnowledgeGraph
from semantica.reasoning.rete_engine import ReteEngine
kg = KnowledgeGraph.load("path/to/graph.db")
engine = ReteEngine(kg)
# Define a risk-flagging rule
engine.add_rule(
condition=lambda n: n.get("risk_score", 0) > 7,
action=lambda n: n.set("flagged", True)
)
engine.run()
print("Flagged nodes:", [n.id for n in kg.nodes if n.get("flagged")])
Summary
ARCHITECTURE.mdprovides the high-level Mermaid diagram and narrative explaining how data flows through all ten architectural layers.semantica/pipeline/pipeline_builder.pyserves as the primary entry point for constructing data pipelines that connect ingestion to storage.- The Ingestion layer is centralized in
semantica/ingest/registry.py, with specific implementations infile_ingestor.pyandweb_ingestor.py. - Knowledge graph construction relies on
semantica/kg/knowledge_graph.pyfor the data model andgraph_builder.pyfor assembly. - Reasoning capabilities are accessed through
semantica/reasoning/reasoner.py, with specific implementations includingdatalog_reasoner.pyandrete_engine.py. - Vector storage is abstracted in
semantica/vector_store/vector_store.py, with FAISS support viafaiss_store.py.
Frequently Asked Questions
What is the entry point for building a data pipeline in Semantica?
The semantica/pipeline/pipeline_builder.py module provides the PipelineBuilder class, which offers a fluent API for declaratively assembling pipelines. This builder connects ingestors, parsers, normalizers, extractors, and knowledge graph constructors into a single executable workflow.
How does Semantica handle vector storage and similarity search?
Semantica abstracts vector storage through semantica/vector_store/vector_store.py, which defines the interface for all backend implementations. The default local implementation uses semantica/vector_store/faiss_store.py for FAISS-based similarity search, while the system also supports Qdrant, Milvus, and other vector databases through the abstract API.
Which files manage conflict detection and deduplication in the knowledge graph?
The semantica/conflicts/__init__.py module identifies contradictory facts before they enter the knowledge graph, while semantica/deduplication/__init__.py removes duplicate information. Both operate during the transition from extraction to knowledge graph construction, ensuring graph integrity before storage in semantica/kg/knowledge_graph.py.
Where is the reasoning logic implemented in Semantica?
Reasoning is coordinated through semantica/reasoning/reasoner.py, which serves as the high-level entry point. Specific implementations include semantica/reasoning/datalog_reasoner.py for Datalog-style inference and semantica/reasoning/rete_engine.py for forward-chaining rule execution over the knowledge graph.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →