How to Build a Semantica Knowledge Graph from Pre-Extracted Entities: A Complete Guide
You can build a Semantica knowledge graph by converting pre-extracted entities into JSON-compatible dictionaries, initializing a Model Context Protocol (MCP) session to access the GraphTool API, and using add_node and add_edge methods to populate the graph before validating with summary() and visualizing with KGVisualizer.
The semantica-agi/semantica repository provides a programmatic, storage-agnostic API for constructing knowledge graphs from any set of pre-extracted entities—whether they originate from LLM outputs, NER pipelines, or manual curation. This guide walks through the canonical workflow using the actual source implementations.
Understanding the Three-Phase Construction Workflow
Building a knowledge graph from pre-extracted entities follows three logical phases, each handled by specific components in the Semantica codebase.
Phase 1: Prepare the Entity Payload
Convert your extracted entities into a list of JSON-compatible dictionaries. Each entity must contain at least an id, a human-readable label, a type (e.g., Person, Organisation, Concept), and optional metadata for provenance tracking.
According to the implementation in semantica_mcp/mcp/tools/graph.py, the add_node method expects these fields to construct graph vertices properly.
Phase 2: Initialize the Graph and Ingest Data
Open a Semantica MCP session and obtain the GraphTool via client.tools.graph(). Use the tool's add_node method (defined at line 17 of semantica_mcp/mcp/tools/graph.py) to insert entities and add_edge (line 125) to create relationships. For bulk ingestion, the SeedManager class (line 101 of semantica/seed/seed_manager.py) can ingest CSV or JSON files directly, handling deduplication and type inference automatically.
Phase 3: Validate and Visualize the Knowledge Graph
After insertion, retrieve a high-level statistical overview using graph.summary() (implemented at line 150 of semantica_mcp/mcp/tools/graph.py). For interactive exploration, instantiate the KGVisualizer class from semantica/visualization/kg_visualizer.py (line 18) to render the graph using Plotly.
Step-by-Step Implementation Guide
Initialize the MCP Client and Graph Tool
Start by establishing a connection to the Semantica MCP server. The SemanticaMCP class in semantica_mcp/mcp/session.py provides an in-process client for local development.
from semantica_mcp.mcp.session import SemanticaMCP
# Initialize client and obtain the graph tool
client = SemanticaMCP()
graph_tool = client.tools.graph()
Format Your Pre-Extracted Entities
Structure your entities to match the expected schema. This example shows entities that might come from an LLM extraction pipeline or SpaCy NER:
entities = [
{"id": "e1", "label": "Alice", "type": "Person", "metadata": {"source": "doc1"}},
{"id": "e2", "label": "Acme Corp", "type": "Company", "metadata": {"source": "doc2"}},
{"id": "e3", "label": "Quantum AI", "type": "Concept", "metadata": {"source": "doc3"}},
]
Add Entities as Nodes
Iterate through your entity list and call add_node for each entry. This method abstracts the underlying storage backend (in-memory, ArangoDB, or Neo4j), keeping your code storage-agnostic.
for ent in entities:
graph_tool.add_node(
node_id=ent["id"],
label=ent["label"],
node_type=ent["type"],
metadata=ent["metadata"],
)
Create Relationships with add_edge
Define relationships between entities using the add_edge method. Each edge requires a source ID, target ID, and relationship label:
relationships = [
{"source": "e1", "target": "e2", "label": "EMPLOYED_BY"},
{"source": "e2", "target": "e3", "label": "WORKS_ON"},
]
for rel in relationships:
graph_tool.add_edge(
source_id=rel["source"],
target_id=rel["target"],
label=rel["label"],
)
Validate Graph Construction
Verify successful ingestion by calling the summary method, which returns node counts, edge counts, and type breakdowns:
summary = graph_tool.summary()
print(summary)
Alternative: Bulk Import with SeedManager
For large-scale ingestion, bypass the explicit Python loop and use the SeedManager to load flat files directly. This approach handles batch processing, deduplication, and metadata preservation automatically.
from semantica.seed.seed_manager import SeedManager
# Initialize with your MCP client
seed_manager = SeedManager(client)
# Load from CSV (columns: id, label, type, metadata)
seed_manager.load_csv("path/to/entities.csv")
# Or load from JSON
seed_manager.load_json("path/to/entities.json")
The SeedManager implementation in semantica/seed/seed_manager.py provides efficient bulk ingestion while maintaining data integrity constraints defined in the graph schema.
Visualizing and Exporting Your Knowledge Graph
Interactive Visualization
Generate interactive network visualizations using the KGVisualizer class. The implementation in semantica/visualization/kg_visualizer.py wraps Plotly and NetworkX for browser-based exploration:
from semantica.visualization.kg_visualizer import KGVisualizer
visualizer = KGVisualizer(client)
visualizer.visualize() # Opens interactive HTML view
For temporal graphs (entities with timestamps), use the TemporalVisualizer in semantica/visualization/temporal_visualizer.py instead.
Export Utilities
Export your constructed graph to standard formats using the export tools defined in semantica_mcp/mcp/tools/export.py:
- JSON for machine-readable interchange
- CSV for tabular analysis
- Parquet for high-performance analytics
- GraphML for compatibility with Gephi and other graph analysis tools
Summary
- Prepare entities as JSON-compatible dictionaries with
id,label,type, and optionalmetadatafields - Initialize the GraphTool via
SemanticaMCP().tools.graph()to access storage-agnostic mutation methods - Add nodes using
add_node()(line 17 ofsemantica_mcp/mcp/tools/graph.py) and edges usingadd_edge()(line 125) - Bulk load large datasets using
SeedManagerfromsemantica/seed/seed_manager.pyfor automatic deduplication - Validate construction with
summary()and visualize withKGVisualizerfromsemantica/visualization/kg_visualizer.py
Frequently Asked Questions
What data format do I need for pre-extracted entities?
Pre-extracted entities must be formatted as JSON-compatible dictionaries containing at minimum an id (unique identifier), label (human-readable name), and type (entity category). Optional metadata fields can store provenance information, confidence scores, or source document references. The add_node method in semantica_mcp/mcp/tools/graph.py accepts these parameters directly.
Can I use the SeedManager for JSON files as well as CSV?
Yes. The SeedManager class in semantica/seed/seed_manager.py provides both load_csv() and load_json() methods. JSON files should contain an array of objects matching the entity schema, while CSV files should include columns for id, label, type, and metadata. The manager automatically handles type inference and entity deduplication during ingestion.
How do I handle entity deduplication during bulk import?
The SeedManager handles deduplication automatically by checking for existing node IDs before insertion. If you are using the low-level GraphTool API directly, you should implement your own deduplication logic by querying existing nodes with search_nodes before calling add_node, or ensure your entity extraction pipeline assigns consistent, deterministic identifiers to avoid duplicates.
Is the GraphTool storage-agnostic?
Yes. The GraphTool implementation in semantica_mcp/mcp/tools/graph.py abstracts the underlying storage layer, supporting in-memory graphs for development, ArangoDB for production document-graph hybrid storage, and Neo4j for native graph persistence. Your construction code remains identical regardless of the backend configuration, enabling seamless migration from prototyping to production scales.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →