Semantica Knowledge Graph Construction: Core Files and Architecture
Semantica constructs knowledge graphs through a three-layer architecture comprising a unified Graph Store core, pluggable database backend implementations, and a semantic extraction pipeline that transforms raw documents into structured nodes and edges.
Semantica is an open-source AGI framework designed for automated knowledge graph construction from unstructured data. The repository semantica-agi/semantica organizes its graph construction capabilities into distinct modules that separate semantic analysis from storage operations, enabling backend portability without modifying extraction logic.
Core Architecture Overview
The knowledge graph construction pipeline in Semantica is organized into three strategic layers. The Graph Store Core provides backend-agnostic CRUD operations and analytics. The Backend Implementations layer contains database-specific drivers for Neo4j, FalkorDB, Amazon Neptune, and Apache Age. The Semantic Extraction layer parses documents and conversations into graph-compatible structures. This separation allows developers to swap from Neo4j to FalkorDB simply by changing a configuration string while preserving the same extraction and ingestion code.
Graph Store Core Files
The central nervous system of Semantica's knowledge graph construction resides in semantica/graph_store/graph_store.py. This file defines the GraphStore class, which serves as the primary interface for all graph operations.
Unified Interface
The GraphStore class exposes high-level methods including add_nodes(), add_edges(), and shortest_path() that function identically across all supported backends. Internally, the class maintains an _app_node_id_map dictionary to reconcile user-provided string identifiers with the internal numeric IDs returned by graph databases. The file also contains specialized manager classes—NodeManager, RelationshipManager, QueryEngine, and GraphAnalytics—that handle specific aspects of graph manipulation.
Backend Registration
The file semantica/graph_store/registry.py implements a plugin architecture that maps backend identifier strings to their concrete implementations. When instantiating a GraphStore with the backend="neo4j" parameter, the registry resolves this to the Neo4j driver class, enabling the core store to remain agnostic of specific database dialects.
Database Backend Implementations
Semantica supports four production-grade graph databases, each implemented in dedicated modules that adhere to a common interface requiring methods like create_node(), create_relationship(), and execute_query().
Neo4j Support
The semantica/graph_store/neo4j_store.py file contains the production Neo4j implementation, including Neo4jDriver, Neo4jSession, and Neo4jTransaction classes. These wrappers handle connection pooling, Cypher query execution, and transaction lifecycle management. The module implements concrete CRUD operations through create_node() and create_relationship() methods that translate generic graph operations into Cypher statements.
Alternative Backends
- FalkorDB: Implemented in
semantica/graph_store/falkordb_store.py, providing RedisGraph compatibility with Redis-backed storage. - Amazon Neptune: The
semantica/graph_store/amazon_neptune.pymodule supports AWS's managed graph service through both Gremlin and OpenCypher query languages. - Apache Age: Located in
semantica/graph_store/age_store.py, this backend enables PostgreSQL-based graph storage using the Apache Age extension.
Semantic Extraction Pipeline
Before data enters the graph store, Semantica's extraction layer transforms raw text into standardized graph elements consisting of node dictionaries (with id, type/labels, and properties) and edge dictionaries (with source_id, target_id, type, and properties).
Semantic Analyzer
The semantica/semantic_extract/semantic_analyzer.py file orchestrates the extraction workflow, coordinating multiple specialized extractors to produce the dual collections required by the Graph Store. This module handles batch processing and ensures consistency between extracted entities and their relationships.
Entity and Relation Extractors
- Triplet Extraction:
semantica/semantic_extract/triplet_extractor.pyextracts (subject, predicate, object) triples from unstructured text using LLM-based or pattern-based approaches. - NER:
semantica/semantic_extract/ner_extractor.pyperforms named entity recognition, identifying entities like persons, organizations, and locations using either spaCy models or large language models. - Network Assembly:
semantica/semantic_extract/semantic_network_extractor.pycombines individual entities and relations into connected graph structures, resolving references and building the final edge list.
Practical Implementation Examples
Initializing the Graph Store
To instantiate a knowledge graph connection with Neo4j:
from semantica.graph_store import GraphStore
store = GraphStore(
backend="neo4j",
uri="bolt://localhost:7687",
user="neo4j",
password="my_secret_password"
)
store.connect()
Ingesting Extracted Data
Load nodes and edges produced by the semantic extraction layer:
nodes = [
{"id": "person_1", "type": "Person", "properties": {"name": "Alice"}},
{"id": "org_1", "type": "Organization", "properties": {"name": "Acme Corp"}}
]
edges = [
{"source_id": "person_1", "target_id": "org_1", "type": "WORKS_FOR", "properties": {"since": 2020}}
]
node_count = store.add_nodes(nodes)
edge_count = store.add_edges(edges)
print(f"Created {node_count} nodes and {edge_count} edges")
Querying and Analytics
Execute graph algorithms without writing backend-specific query languages:
path = store.shortest_path(
start_node_id="person_1",
end_node_id="org_1",
rel_type="WORKS_FOR",
max_depth=5
)
print("Shortest path:", path)
Switching Backends
Change storage backends without modifying extraction or ingestion code:
store = GraphStore(
backend="falkordb",
uri="redis://localhost:6379",
password="redis_pass"
)
store.connect()
# add_nodes() and add_edges() function identically
Utility and Safety Components
Several utility modules ensure reliable knowledge graph construction:
semantica/graph_store/query_sanitize.py: Sanitizes labels, relationship types, and property identifiers before embedding them in Cypher or OpenCypher queries, preventing injection vulnerabilities.semantica/utils/logging.py: Provides centralized logger creation used throughout the graph pipeline for debugging and audit trails.semantica/utils/progress_tracker.py: Reports ingestion progress during long-running graph builds involving millions of nodes.semantica/utils/exceptions.py: Defines custom exception classes for handling backend connection failures, query syntax errors, and validation issues.
Summary
- Semantica knowledge graph construction centers on
semantica/graph_store/graph_store.py, which provides the unifiedGraphStoreinterface supporting Neo4j, FalkorDB, Amazon Neptune, and Apache Age backends. - The backend registry in
registry.pyenables plug-and-play database substitution without code changes to the extraction or ingestion layers. - Semantic extraction files including
semantic_analyzer.py,triplet_extractor.py, andner_extractor.pytransform documents into standardized node and edge dictionaries. - ID reconciliation occurs through the
_app_node_id_mapinternal mapping, bridging user string IDs with backend numeric identifiers. - Safety mechanisms in
query_sanitize.pyensure all dynamically constructed queries are properly escaped before execution.
Frequently Asked Questions
What is the primary file for Semantica's GraphStore class?
The GraphStore class is defined in semantica/graph_store/graph_store.py. This file contains the core interface for knowledge graph construction, including the NodeManager and RelationshipManager classes that implement add_nodes() and add_edges() methods across all supported backends.
Which files handle entity extraction in Semantica?
Entity and relation extraction is managed by files in the semantica/semantic_extract/ directory. The semantic_analyzer.py orchestrates the pipeline, while triplet_extractor.py handles relation extraction, ner_extractor.py performs named entity recognition, and semantic_network_extractor.py assembles the final graph structure from extracted components.
How does Semantica support multiple graph database backends?
Semantica uses a registry pattern in semantica/graph_store/registry.py to map backend strings to concrete implementations. Each backend module (such as neo4j_store.py or falkordb_store.py) implements standardized methods like create_node() and execute_query(), allowing the core GraphStore to operate without database-specific logic.
How does Semantica prevent query injection attacks?
The framework prevents injection through semantica/graph_store/query_sanitize.py, which escapes all labels, relationship types, and identifiers before embedding them in Cypher or OpenCypher strings. This ensures that user-provided data cannot alter query structure or execute arbitrary commands regardless of the backend database in use.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →