# Semantica Knowledge Graph Construction: Core Files and Architecture

> Explore Semantica's knowledge graph construction architecture. Understand the core files and three-layer system for transforming documents into structured data.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: architecture
- Published: 2026-09-10

---

**Semantica constructs knowledge graphs through a three-layer architecture comprising a unified Graph Store core, pluggable database backend implementations, and a semantic extraction pipeline that transforms raw documents into structured nodes and edges.**

Semantica is an open-source AGI framework designed for automated knowledge graph construction from unstructured data. The repository `semantica-agi/semantica` organizes its graph construction capabilities into distinct modules that separate semantic analysis from storage operations, enabling backend portability without modifying extraction logic.

## Core Architecture Overview

The knowledge graph construction pipeline in Semantica is organized into three strategic layers. The **Graph Store Core** provides backend-agnostic CRUD operations and analytics. The **Backend Implementations** layer contains database-specific drivers for Neo4j, FalkorDB, Amazon Neptune, and Apache Age. The **Semantic Extraction** layer parses documents and conversations into graph-compatible structures. This separation allows developers to swap from Neo4j to FalkorDB simply by changing a configuration string while preserving the same extraction and ingestion code.

## Graph Store Core Files

The central nervous system of Semantica's knowledge graph construction resides in [`semantica/graph_store/graph_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/graph_store.py). This file defines the `GraphStore` class, which serves as the primary interface for all graph operations.

### Unified Interface

The `GraphStore` class exposes high-level methods including `add_nodes()`, `add_edges()`, and `shortest_path()` that function identically across all supported backends. Internally, the class maintains an `_app_node_id_map` dictionary to reconcile user-provided string identifiers with the internal numeric IDs returned by graph databases. The file also contains specialized manager classes—`NodeManager`, `RelationshipManager`, `QueryEngine`, and `GraphAnalytics`—that handle specific aspects of graph manipulation.

### Backend Registration

The file [`semantica/graph_store/registry.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/registry.py) implements a plugin architecture that maps backend identifier strings to their concrete implementations. When instantiating a `GraphStore` with the `backend="neo4j"` parameter, the registry resolves this to the Neo4j driver class, enabling the core store to remain agnostic of specific database dialects.

## Database Backend Implementations

Semantica supports four production-grade graph databases, each implemented in dedicated modules that adhere to a common interface requiring methods like `create_node()`, `create_relationship()`, and `execute_query()`.

### Neo4j Support

The [`semantica/graph_store/neo4j_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/neo4j_store.py) file contains the production Neo4j implementation, including `Neo4jDriver`, `Neo4jSession`, and `Neo4jTransaction` classes. These wrappers handle connection pooling, Cypher query execution, and transaction lifecycle management. The module implements concrete CRUD operations through `create_node()` and `create_relationship()` methods that translate generic graph operations into Cypher statements.

### Alternative Backends

- **FalkorDB**: Implemented in [`semantica/graph_store/falkordb_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/falkordb_store.py), providing RedisGraph compatibility with Redis-backed storage.
- **Amazon Neptune**: The [`semantica/graph_store/amazon_neptune.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/amazon_neptune.py) module supports AWS's managed graph service through both Gremlin and OpenCypher query languages.
- **Apache Age**: Located in [`semantica/graph_store/age_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/age_store.py), this backend enables PostgreSQL-based graph storage using the Apache Age extension.

## Semantic Extraction Pipeline

Before data enters the graph store, Semantica's extraction layer transforms raw text into standardized graph elements consisting of node dictionaries (with `id`, `type`/`labels`, and `properties`) and edge dictionaries (with `source_id`, `target_id`, `type`, and `properties`).

### Semantic Analyzer

The [`semantica/semantic_extract/semantic_analyzer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/semantic_analyzer.py) file orchestrates the extraction workflow, coordinating multiple specialized extractors to produce the dual collections required by the Graph Store. This module handles batch processing and ensures consistency between extracted entities and their relationships.

### Entity and Relation Extractors

- **Triplet Extraction**: [`semantica/semantic_extract/triplet_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/triplet_extractor.py) extracts (subject, predicate, object) triples from unstructured text using LLM-based or pattern-based approaches.
- **NER**: [`semantica/semantic_extract/ner_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/ner_extractor.py) performs named entity recognition, identifying entities like persons, organizations, and locations using either spaCy models or large language models.
- **Network Assembly**: [`semantica/semantic_extract/semantic_network_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/semantic_network_extractor.py) combines individual entities and relations into connected graph structures, resolving references and building the final edge list.

## Practical Implementation Examples

### Initializing the Graph Store

To instantiate a knowledge graph connection with Neo4j:

```python
from semantica.graph_store import GraphStore

store = GraphStore(
    backend="neo4j",
    uri="bolt://localhost:7687",
    user="neo4j",
    password="my_secret_password"
)
store.connect()

```

### Ingesting Extracted Data

Load nodes and edges produced by the semantic extraction layer:

```python
nodes = [
    {"id": "person_1", "type": "Person", "properties": {"name": "Alice"}},
    {"id": "org_1", "type": "Organization", "properties": {"name": "Acme Corp"}}
]

edges = [
    {"source_id": "person_1", "target_id": "org_1", "type": "WORKS_FOR", "properties": {"since": 2020}}
]

node_count = store.add_nodes(nodes)
edge_count = store.add_edges(edges)
print(f"Created {node_count} nodes and {edge_count} edges")

```

### Querying and Analytics

Execute graph algorithms without writing backend-specific query languages:

```python
path = store.shortest_path(
    start_node_id="person_1",
    end_node_id="org_1",
    rel_type="WORKS_FOR",
    max_depth=5
)
print("Shortest path:", path)

```

### Switching Backends

Change storage backends without modifying extraction or ingestion code:

```python
store = GraphStore(
    backend="falkordb",
    uri="redis://localhost:6379",
    password="redis_pass"
)
store.connect()

# add_nodes() and add_edges() function identically

```

## Utility and Safety Components

Several utility modules ensure reliable knowledge graph construction:

- **[`semantica/graph_store/query_sanitize.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/query_sanitize.py)**: Sanitizes labels, relationship types, and property identifiers before embedding them in Cypher or OpenCypher queries, preventing injection vulnerabilities.
- **[`semantica/utils/logging.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/utils/logging.py)**: Provides centralized logger creation used throughout the graph pipeline for debugging and audit trails.
- **[`semantica/utils/progress_tracker.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/utils/progress_tracker.py)**: Reports ingestion progress during long-running graph builds involving millions of nodes.
- **[`semantica/utils/exceptions.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/utils/exceptions.py)**: Defines custom exception classes for handling backend connection failures, query syntax errors, and validation issues.

## Summary

- **Semantica knowledge graph construction** centers on [`semantica/graph_store/graph_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/graph_store.py), which provides the unified `GraphStore` interface supporting Neo4j, FalkorDB, Amazon Neptune, and Apache Age backends.
- The **backend registry** in [`registry.py`](https://github.com/semantica-agi/semantica/blob/main/registry.py) enables plug-and-play database substitution without code changes to the extraction or ingestion layers.
- **Semantic extraction** files including [`semantic_analyzer.py`](https://github.com/semantica-agi/semantica/blob/main/semantic_analyzer.py), [`triplet_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/triplet_extractor.py), and [`ner_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/ner_extractor.py) transform documents into standardized node and edge dictionaries.
- **ID reconciliation** occurs through the `_app_node_id_map` internal mapping, bridging user string IDs with backend numeric identifiers.
- **Safety mechanisms** in [`query_sanitize.py`](https://github.com/semantica-agi/semantica/blob/main/query_sanitize.py) ensure all dynamically constructed queries are properly escaped before execution.

## Frequently Asked Questions

### What is the primary file for Semantica's GraphStore class?

The `GraphStore` class is defined in [`semantica/graph_store/graph_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/graph_store.py). This file contains the core interface for knowledge graph construction, including the `NodeManager` and `RelationshipManager` classes that implement `add_nodes()` and `add_edges()` methods across all supported backends.

### Which files handle entity extraction in Semantica?

Entity and relation extraction is managed by files in the `semantica/semantic_extract/` directory. The [`semantic_analyzer.py`](https://github.com/semantica-agi/semantica/blob/main/semantic_analyzer.py) orchestrates the pipeline, while [`triplet_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/triplet_extractor.py) handles relation extraction, [`ner_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/ner_extractor.py) performs named entity recognition, and [`semantic_network_extractor.py`](https://github.com/semantica-agi/semantica/blob/main/semantic_network_extractor.py) assembles the final graph structure from extracted components.

### How does Semantica support multiple graph database backends?

Semantica uses a registry pattern in [`semantica/graph_store/registry.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/registry.py) to map backend strings to concrete implementations. Each backend module (such as [`neo4j_store.py`](https://github.com/semantica-agi/semantica/blob/main/neo4j_store.py) or [`falkordb_store.py`](https://github.com/semantica-agi/semantica/blob/main/falkordb_store.py)) implements standardized methods like `create_node()` and `execute_query()`, allowing the core `GraphStore` to operate without database-specific logic.

### How does Semantica prevent query injection attacks?

The framework prevents injection through [`semantica/graph_store/query_sanitize.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/query_sanitize.py), which escapes all labels, relationship types, and identifiers before embedding them in Cypher or OpenCypher strings. This ensures that user-provided data cannot alter query structure or execute arbitrary commands regardless of the backend database in use.