Semantica Knowledge Graph Construction: Core Files and Architecture

Semantica constructs knowledge graphs through a three-layer architecture comprising a unified Graph Store core, pluggable database backend implementations, and a semantic extraction pipeline that transforms raw documents into structured nodes and edges.

Semantica is an open-source AGI framework designed for automated knowledge graph construction from unstructured data. The repository semantica-agi/semantica organizes its graph construction capabilities into distinct modules that separate semantic analysis from storage operations, enabling backend portability without modifying extraction logic.

Core Architecture Overview

The knowledge graph construction pipeline in Semantica is organized into three strategic layers. The Graph Store Core provides backend-agnostic CRUD operations and analytics. The Backend Implementations layer contains database-specific drivers for Neo4j, FalkorDB, Amazon Neptune, and Apache Age. The Semantic Extraction layer parses documents and conversations into graph-compatible structures. This separation allows developers to swap from Neo4j to FalkorDB simply by changing a configuration string while preserving the same extraction and ingestion code.

Graph Store Core Files

The central nervous system of Semantica's knowledge graph construction resides in semantica/graph_store/graph_store.py. This file defines the GraphStore class, which serves as the primary interface for all graph operations.

Unified Interface

The GraphStore class exposes high-level methods including add_nodes(), add_edges(), and shortest_path() that function identically across all supported backends. Internally, the class maintains an _app_node_id_map dictionary to reconcile user-provided string identifiers with the internal numeric IDs returned by graph databases. The file also contains specialized manager classes—NodeManager, RelationshipManager, QueryEngine, and GraphAnalytics—that handle specific aspects of graph manipulation.

Backend Registration

The file semantica/graph_store/registry.py implements a plugin architecture that maps backend identifier strings to their concrete implementations. When instantiating a GraphStore with the backend="neo4j" parameter, the registry resolves this to the Neo4j driver class, enabling the core store to remain agnostic of specific database dialects.

Database Backend Implementations

Semantica supports four production-grade graph databases, each implemented in dedicated modules that adhere to a common interface requiring methods like create_node(), create_relationship(), and execute_query().

Neo4j Support

The semantica/graph_store/neo4j_store.py file contains the production Neo4j implementation, including Neo4jDriver, Neo4jSession, and Neo4jTransaction classes. These wrappers handle connection pooling, Cypher query execution, and transaction lifecycle management. The module implements concrete CRUD operations through create_node() and create_relationship() methods that translate generic graph operations into Cypher statements.

Alternative Backends

Semantic Extraction Pipeline

Before data enters the graph store, Semantica's extraction layer transforms raw text into standardized graph elements consisting of node dictionaries (with id, type/labels, and properties) and edge dictionaries (with source_id, target_id, type, and properties).

Semantic Analyzer

The semantica/semantic_extract/semantic_analyzer.py file orchestrates the extraction workflow, coordinating multiple specialized extractors to produce the dual collections required by the Graph Store. This module handles batch processing and ensures consistency between extracted entities and their relationships.

Entity and Relation Extractors

Practical Implementation Examples

Initializing the Graph Store

To instantiate a knowledge graph connection with Neo4j:

from semantica.graph_store import GraphStore

store = GraphStore(
    backend="neo4j",
    uri="bolt://localhost:7687",
    user="neo4j",
    password="my_secret_password"
)
store.connect()

Ingesting Extracted Data

Load nodes and edges produced by the semantic extraction layer:

nodes = [
    {"id": "person_1", "type": "Person", "properties": {"name": "Alice"}},
    {"id": "org_1", "type": "Organization", "properties": {"name": "Acme Corp"}}
]

edges = [
    {"source_id": "person_1", "target_id": "org_1", "type": "WORKS_FOR", "properties": {"since": 2020}}
]

node_count = store.add_nodes(nodes)
edge_count = store.add_edges(edges)
print(f"Created {node_count} nodes and {edge_count} edges")

Querying and Analytics

Execute graph algorithms without writing backend-specific query languages:

path = store.shortest_path(
    start_node_id="person_1",
    end_node_id="org_1",
    rel_type="WORKS_FOR",
    max_depth=5
)
print("Shortest path:", path)

Switching Backends

Change storage backends without modifying extraction or ingestion code:

store = GraphStore(
    backend="falkordb",
    uri="redis://localhost:6379",
    password="redis_pass"
)
store.connect()

# add_nodes() and add_edges() function identically

Utility and Safety Components

Several utility modules ensure reliable knowledge graph construction:

Summary

  • Semantica knowledge graph construction centers on semantica/graph_store/graph_store.py, which provides the unified GraphStore interface supporting Neo4j, FalkorDB, Amazon Neptune, and Apache Age backends.
  • The backend registry in registry.py enables plug-and-play database substitution without code changes to the extraction or ingestion layers.
  • Semantic extraction files including semantic_analyzer.py, triplet_extractor.py, and ner_extractor.py transform documents into standardized node and edge dictionaries.
  • ID reconciliation occurs through the _app_node_id_map internal mapping, bridging user string IDs with backend numeric identifiers.
  • Safety mechanisms in query_sanitize.py ensure all dynamically constructed queries are properly escaped before execution.

Frequently Asked Questions

What is the primary file for Semantica's GraphStore class?

The GraphStore class is defined in semantica/graph_store/graph_store.py. This file contains the core interface for knowledge graph construction, including the NodeManager and RelationshipManager classes that implement add_nodes() and add_edges() methods across all supported backends.

Which files handle entity extraction in Semantica?

Entity and relation extraction is managed by files in the semantica/semantic_extract/ directory. The semantic_analyzer.py orchestrates the pipeline, while triplet_extractor.py handles relation extraction, ner_extractor.py performs named entity recognition, and semantic_network_extractor.py assembles the final graph structure from extracted components.

How does Semantica support multiple graph database backends?

Semantica uses a registry pattern in semantica/graph_store/registry.py to map backend strings to concrete implementations. Each backend module (such as neo4j_store.py or falkordb_store.py) implements standardized methods like create_node() and execute_query(), allowing the core GraphStore to operate without database-specific logic.

How does Semantica prevent query injection attacks?

The framework prevents injection through semantica/graph_store/query_sanitize.py, which escapes all labels, relationship types, and identifiers before embedding them in Cypher or OpenCypher strings. This ensures that user-provided data cannot alter query structure or execute arbitrary commands regardless of the backend database in use.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →