# What Are the Core Components of Semantica’s Architecture?

> Explore Semantica's architecture, a layered data-to-knowledge pipeline. Discover its 20+ Python packages for ingesting data, extracting triples, and building a queryable knowledge graph with provenance.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: architecture
- Published: 2026-09-09

---

**Semantica’s architecture comprises a layered data-to-knowledge pipeline with 20+ specialized Python packages that ingest raw data, parse and normalize content, extract semantic triples, resolve conflicts, and construct a queryable knowledge graph with full provenance tracking.**

Semantica (semantica-agi/semantica) is an open-source decision intelligence platform that transforms unstructured data into actionable knowledge graphs. The codebase implements a modular, pipeline-based architecture where each layer handles a specific transformation step. Understanding these core components is essential for customizing ingestion workflows or extending the platform’s reasoning capabilities.

## Data Ingestion and Source Connectors

The **Sources** layer (`semantica.ingest`) provides connectors for files, web pages, databases, cloud stores, and streams. Concrete implementations like `FileIngestor` in [`semantica/ingest/file_ingestor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/file_ingestor.py) and `WebIngestor` transform source-specific payloads into a unified raw document representation. This abstraction allows the pipeline to consume everything from local PDFs to real-time API streams through a consistent interface.

## Document Processing Pipeline

Once ingested, documents flow through three processing stages located in `semantica.parse`, `semantica.normalize`, and `semantica.split`.

### Parsing and Structuring

The `DocumentParser` class in [`semantica/parse/document_parser.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/parse/document_parser.py) converts raw bytes into structured objects, extracting text, tables, code blocks, and images. This component handles format detection and delegates to specialized parsers for PDFs, Word documents, HTML, and markdown.

### Normalization

The **Normalize** layer (`semantica.normalize`) cleans and canonicalizes content. The `TextNormalizer` in [`semantica/normalize/text_normalizer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/normalize/text_normalizer.py) standardizes dates, numbers, and entities, ensuring consistent representation across sources before semantic extraction begins.

### Chunking and Splitting

The **Split** layer breaks documents into logical, context-aware chunks. The `GraphBasedSplitter` in [`semantica/split/graph_based.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/split/graph_based.py) uses semantic relationships to create entity-aware segments, preserving context boundaries better than naive character-count approaches.

## Semantic Extraction and Data Quality

After preprocessing, the **Extract** layer (`semantica.semantic_extract`) performs named-entity recognition, relation extraction, and triplet generation. The `TripletExtractor` in [`semantica/semantic_extract/triple_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/triple_extractor.py) identifies subject-predicate-object relationships.

### Conflict Detection

The `ConflictDetector` in [`semantica/conflicts/conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/conflict_detector.py) scans extracted triples for contradictory statements across sources. This component flags inconsistencies before they enter the knowledge graph, maintaining data integrity.

### Deduplication

The `DuplicateDetector` in [`semantica/deduplication/duplicate_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/deduplication/duplicate_detector.py) merges redundant entities and prunes duplicate relationships. This deduplication step reduces graph noise and improves query performance.

## Knowledge Graph Construction and Reasoning

The construction phase assembles cleaned triples into a formal knowledge graph.

### Graph Assembly

The `GraphBuilder` in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py) constructs the internal KG data model, connecting entities via typed relationships and preparing the graph for persistence.

### Ontology Management

The **Ontology** layer (`semantica.ontology`) generates and validates domain schemas. The `OntologyGenerator` in [`semantica/ontology/generator.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ontology/generator.py) produces OWL, SHACL, and SKOS definitions, ensuring the KG adheres to semantic standards.

### Reasoning Engine

The **Reasoning** layer (`semantica.reasoning`) applies inference rules to derive implicit knowledge. The `SPARQLReasoner` in [`semantica/reasoning/sparql_reasoner.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/reasoning/sparql_reasoner.py) executes Datalog and SPARQL queries to enrich the graph with logical consequences and transitive relationships.

## Storage, Provenance, and Decision Intelligence

Semantica implements the **Decision Intelligence Lifecycle** (record → link → query → govern → audit) through specialized storage and tracking components.

### Provenance Tracking

Every artifact carries W3C PROV-O metadata managed by the `ProvenanceManager` in [`semantica/provenance/provenance_manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/provenance_manager.py). This creates an immutable audit trail linking every extracted triple back to its original source document and extraction timestamp.

### Decision Context

The **Context** layer (`semantica.context`) builds causal decision graphs. The `DecisionRecorder` in [`semantica/context/decision_recorder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/context/decision_recorder.py) records policy evaluations and agent decisions, enabling explainable AI workflows.

### Vector and Graph Storage

For retrieval and persistence, Semantica uses dual storage backends. The `VectorStore` in [`semantica/vector_store/vector_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/vector_store/vector_store.py) manages embeddings for similarity search (supporting FAISS, Qdrant, and Weaviate), while `Neo4jStore` in [`semantica/graph_store/neo4j_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/neo4j_store.py) persists the knowledge graph to graph databases.

## Export, Visualization, and Services

The final layers expose the enriched knowledge graph to downstream applications.

### Export Formats

The **Export** layer (`semantica.export`) serializes graphs to RDF, JSON-LD, Parquet, CSV, and GraphML. The `ParquetExporter` in [`semantica/export/parquet_exporter.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/export/parquet_exporter.py) enables efficient analytics workflows with columnar compression.

### Visualization

The `SemanticNetworkVisualizer` in [`semantica/visualization/semantic_network_visualizer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/semantic_network_visualizer.py) renders interactive embeddings, temporal graphs, and ontology hierarchies for exploratory analysis.

### Service Interfaces

The architecture exposes functionality through multiple interfaces defined in [`semantica/cli.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/cli.py), including a REST API, command-line tools, and an MCP Server for integration with the Knowledge Explorer UI.

## End-to-End Integration Example

The following Python workflow demonstrates how these components connect to process a document into an enriched knowledge graph:

```python

# 1️⃣ Ingest a local PDF file

from semantica.ingest.file_ingestor import FileIngestor
raw_doc = FileIngestor().ingest("sample.pdf")

# 2️⃣ Parse the raw document

from semantica.parse.document_parser import DocumentParser
parsed = DocumentParser().parse(raw_doc)

# 3️⃣ Normalize the parsed output

from semantica.normalize.text_normalizer import TextNormalizer
norm = TextNormalizer().normalize(parsed)

# 4️⃣ Split into logical chunks

from semantica.split.graph_based import GraphBasedSplitter
chunks = GraphBasedSplitter().split(norm)

# 5️⃣ Extract entities & relationships

from semantica.semantic_extract.triple_extractor import TripletExtractor
triples = TripletExtractor().extract(chunks)

# 6️⃣ Resolve conflicts and deduplicate

from semantica.conflicts.conflict_resolver import ConflictResolver
clean_triples = ConflictResolver().resolve(triples)

from semantica.deduplication.entity_merger import EntityMerger
deduped = EntityMerger().merge(clean_triples)

# 7️⃣ Build the knowledge graph

from semantica.kg.graph_builder import GraphBuilder
kg = GraphBuilder().build(deduped)

# 8️⃣ Enrich with ontology and reasoning

from semantica.ontology.generator import OntologyGenerator
ont = OntologyGenerator().generate(kg)

from semantica.reasoning.sparql_reasoner import SPARQLReasoner
enriched_kg = SPARQLReasoner().apply(kg, ont)

# 9️⃣ Store embeddings for similarity search

from semantica.vector_store.faiss_store import FAISSStore
vec_store = FAISSStore()
vec_store.index(enriched_kg)

# 🔟 Export the enriched KG to Parquet

from semantica.export.parquet_exporter import ParquetExporter
ParquetExporter(compression="snappy").export_knowledge_graph(
    {"entities": enriched_kg.entities, "relationships": enriched_kg.relationships},
    Path("output/knowledge_graph")
)

```

This snippet mirrors the data flow shown in the repository’s [`ARCHITECTURE.md`](https://github.com/semantica-agi/semantica/blob/main/ARCHITECTURE.md), connecting each core component via its public API.

## Summary

- **Semantica implements a 20-layer pipeline** organized into ingestion, processing, extraction, construction, and service tiers.
- **Each layer is a self-contained Python package** (e.g., `semantica.ingest`, `semantica.semantic_extract`) with specific responsibilities.
- **Data quality is enforced** through explicit conflict detection ([`conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/conflict_detector.py)) and deduplication ([`duplicate_detector.py`](https://github.com/semantica-agi/semantica/blob/main/duplicate_detector.py)) stages.
- **Full provenance tracking** via W3C PROV-O ensures every triple maintains an auditable link to its source.
- **Dual storage backends** separate vector embeddings for similarity search from graph structures for relationship queries.
- **Multiple export formats** (RDF, Parquet, JSON-LD) and service interfaces (CLI, REST, MCP) enable integration with existing data stacks.

## Frequently Asked Questions

### How does Semantica handle data ingestion from different sources?

Semantica abstracts source complexity through the `semantica.ingest` package. Concrete implementations like `FileIngestor` ([`semantica/ingest/file_ingestor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/file_ingestor.py)) and `WebIngestor` normalize diverse inputs (PDFs, databases, APIs) into a unified raw document format, allowing downstream components to process heterogeneous data through a single interface.

### What is the role of the conflict detection layer in Semantica's architecture?

The conflict detection layer ([`semantica/conflicts/conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/conflict_detector.py)) scans extracted semantic triples for contradictory statements across multiple sources. By identifying inconsistencies before they enter the knowledge graph, this component maintains data integrity and prevents logical contradictions from affecting reasoning results.

### How does Semantica track provenance throughout the knowledge graph pipeline?

The `ProvenanceManager` in [`semantica/provenance/provenance_manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/provenance/provenance_manager.py) implements W3C PROV-O standards to record the origin, transformation history, and timestamps for every artifact. This creates an immutable audit trail that links each extracted entity and relationship back to its original source document and processing step.

### Can Semantica export knowledge graphs to standard semantic web formats?

Yes, the `semantica.export` layer supports multiple serialization formats including RDF, JSON-LD, and GraphML for semantic web integration, as well as Parquet and CSV for analytics workflows. The `ParquetExporter` in [`semantica/export/parquet_exporter.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/export/parquet_exporter.py) enables compressed, columnar storage suitable for large-scale data science pipelines.