What Are the Core Components of Semantica’s Architecture?

Semantica’s architecture comprises a layered data-to-knowledge pipeline with 20+ specialized Python packages that ingest raw data, parse and normalize content, extract semantic triples, resolve conflicts, and construct a queryable knowledge graph with full provenance tracking.

Semantica (semantica-agi/semantica) is an open-source decision intelligence platform that transforms unstructured data into actionable knowledge graphs. The codebase implements a modular, pipeline-based architecture where each layer handles a specific transformation step. Understanding these core components is essential for customizing ingestion workflows or extending the platform’s reasoning capabilities.

Data Ingestion and Source Connectors

The Sources layer (semantica.ingest) provides connectors for files, web pages, databases, cloud stores, and streams. Concrete implementations like FileIngestor in semantica/ingest/file_ingestor.py and WebIngestor transform source-specific payloads into a unified raw document representation. This abstraction allows the pipeline to consume everything from local PDFs to real-time API streams through a consistent interface.

Document Processing Pipeline

Once ingested, documents flow through three processing stages located in semantica.parse, semantica.normalize, and semantica.split.

Parsing and Structuring

The DocumentParser class in semantica/parse/document_parser.py converts raw bytes into structured objects, extracting text, tables, code blocks, and images. This component handles format detection and delegates to specialized parsers for PDFs, Word documents, HTML, and markdown.

Normalization

The Normalize layer (semantica.normalize) cleans and canonicalizes content. The TextNormalizer in semantica/normalize/text_normalizer.py standardizes dates, numbers, and entities, ensuring consistent representation across sources before semantic extraction begins.

Chunking and Splitting

The Split layer breaks documents into logical, context-aware chunks. The GraphBasedSplitter in semantica/split/graph_based.py uses semantic relationships to create entity-aware segments, preserving context boundaries better than naive character-count approaches.

Semantic Extraction and Data Quality

After preprocessing, the Extract layer (semantica.semantic_extract) performs named-entity recognition, relation extraction, and triplet generation. The TripletExtractor in semantica/semantic_extract/triple_extractor.py identifies subject-predicate-object relationships.

Conflict Detection

The ConflictDetector in semantica/conflicts/conflict_detector.py scans extracted triples for contradictory statements across sources. This component flags inconsistencies before they enter the knowledge graph, maintaining data integrity.

Deduplication

The DuplicateDetector in semantica/deduplication/duplicate_detector.py merges redundant entities and prunes duplicate relationships. This deduplication step reduces graph noise and improves query performance.

Knowledge Graph Construction and Reasoning

The construction phase assembles cleaned triples into a formal knowledge graph.

Graph Assembly

The GraphBuilder in semantica/kg/graph_builder.py constructs the internal KG data model, connecting entities via typed relationships and preparing the graph for persistence.

Ontology Management

The Ontology layer (semantica.ontology) generates and validates domain schemas. The OntologyGenerator in semantica/ontology/generator.py produces OWL, SHACL, and SKOS definitions, ensuring the KG adheres to semantic standards.

Reasoning Engine

The Reasoning layer (semantica.reasoning) applies inference rules to derive implicit knowledge. The SPARQLReasoner in semantica/reasoning/sparql_reasoner.py executes Datalog and SPARQL queries to enrich the graph with logical consequences and transitive relationships.

Storage, Provenance, and Decision Intelligence

Semantica implements the Decision Intelligence Lifecycle (record → link → query → govern → audit) through specialized storage and tracking components.

Provenance Tracking

Every artifact carries W3C PROV-O metadata managed by the ProvenanceManager in semantica/provenance/provenance_manager.py. This creates an immutable audit trail linking every extracted triple back to its original source document and extraction timestamp.

Decision Context

The Context layer (semantica.context) builds causal decision graphs. The DecisionRecorder in semantica/context/decision_recorder.py records policy evaluations and agent decisions, enabling explainable AI workflows.

Vector and Graph Storage

For retrieval and persistence, Semantica uses dual storage backends. The VectorStore in semantica/vector_store/vector_store.py manages embeddings for similarity search (supporting FAISS, Qdrant, and Weaviate), while Neo4jStore in semantica/graph_store/neo4j_store.py persists the knowledge graph to graph databases.

Export, Visualization, and Services

The final layers expose the enriched knowledge graph to downstream applications.

Export Formats

The Export layer (semantica.export) serializes graphs to RDF, JSON-LD, Parquet, CSV, and GraphML. The ParquetExporter in semantica/export/parquet_exporter.py enables efficient analytics workflows with columnar compression.

Visualization

The SemanticNetworkVisualizer in semantica/visualization/semantic_network_visualizer.py renders interactive embeddings, temporal graphs, and ontology hierarchies for exploratory analysis.

Service Interfaces

The architecture exposes functionality through multiple interfaces defined in semantica/cli.py, including a REST API, command-line tools, and an MCP Server for integration with the Knowledge Explorer UI.

End-to-End Integration Example

The following Python workflow demonstrates how these components connect to process a document into an enriched knowledge graph:


# 1️⃣ Ingest a local PDF file

from semantica.ingest.file_ingestor import FileIngestor
raw_doc = FileIngestor().ingest("sample.pdf")

# 2️⃣ Parse the raw document

from semantica.parse.document_parser import DocumentParser
parsed = DocumentParser().parse(raw_doc)

# 3️⃣ Normalize the parsed output

from semantica.normalize.text_normalizer import TextNormalizer
norm = TextNormalizer().normalize(parsed)

# 4️⃣ Split into logical chunks

from semantica.split.graph_based import GraphBasedSplitter
chunks = GraphBasedSplitter().split(norm)

# 5️⃣ Extract entities & relationships

from semantica.semantic_extract.triple_extractor import TripletExtractor
triples = TripletExtractor().extract(chunks)

# 6️⃣ Resolve conflicts and deduplicate

from semantica.conflicts.conflict_resolver import ConflictResolver
clean_triples = ConflictResolver().resolve(triples)

from semantica.deduplication.entity_merger import EntityMerger
deduped = EntityMerger().merge(clean_triples)

# 7️⃣ Build the knowledge graph

from semantica.kg.graph_builder import GraphBuilder
kg = GraphBuilder().build(deduped)

# 8️⃣ Enrich with ontology and reasoning

from semantica.ontology.generator import OntologyGenerator
ont = OntologyGenerator().generate(kg)

from semantica.reasoning.sparql_reasoner import SPARQLReasoner
enriched_kg = SPARQLReasoner().apply(kg, ont)

# 9️⃣ Store embeddings for similarity search

from semantica.vector_store.faiss_store import FAISSStore
vec_store = FAISSStore()
vec_store.index(enriched_kg)

# 🔟 Export the enriched KG to Parquet

from semantica.export.parquet_exporter import ParquetExporter
ParquetExporter(compression="snappy").export_knowledge_graph(
    {"entities": enriched_kg.entities, "relationships": enriched_kg.relationships},
    Path("output/knowledge_graph")
)

This snippet mirrors the data flow shown in the repository’s ARCHITECTURE.md, connecting each core component via its public API.

Summary

  • Semantica implements a 20-layer pipeline organized into ingestion, processing, extraction, construction, and service tiers.
  • Each layer is a self-contained Python package (e.g., semantica.ingest, semantica.semantic_extract) with specific responsibilities.
  • Data quality is enforced through explicit conflict detection (conflict_detector.py) and deduplication (duplicate_detector.py) stages.
  • Full provenance tracking via W3C PROV-O ensures every triple maintains an auditable link to its source.
  • Dual storage backends separate vector embeddings for similarity search from graph structures for relationship queries.
  • Multiple export formats (RDF, Parquet, JSON-LD) and service interfaces (CLI, REST, MCP) enable integration with existing data stacks.

Frequently Asked Questions

How does Semantica handle data ingestion from different sources?

Semantica abstracts source complexity through the semantica.ingest package. Concrete implementations like FileIngestor (semantica/ingest/file_ingestor.py) and WebIngestor normalize diverse inputs (PDFs, databases, APIs) into a unified raw document format, allowing downstream components to process heterogeneous data through a single interface.

What is the role of the conflict detection layer in Semantica's architecture?

The conflict detection layer (semantica/conflicts/conflict_detector.py) scans extracted semantic triples for contradictory statements across multiple sources. By identifying inconsistencies before they enter the knowledge graph, this component maintains data integrity and prevents logical contradictions from affecting reasoning results.

How does Semantica track provenance throughout the knowledge graph pipeline?

The ProvenanceManager in semantica/provenance/provenance_manager.py implements W3C PROV-O standards to record the origin, transformation history, and timestamps for every artifact. This creates an immutable audit trail that links each extracted entity and relationship back to its original source document and processing step.

Can Semantica export knowledge graphs to standard semantic web formats?

Yes, the semantica.export layer supports multiple serialization formats including RDF, JSON-LD, and GraphML for semantic web integration, as well as Parquet and CSV for analytics workflows. The ParquetExporter in semantica/export/parquet_exporter.py enables compressed, columnar storage suitable for large-scale data science pipelines.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →