# How to Build a Semantica Knowledge Graph from Pre-Extracted Entities: A Complete Guide

> Learn to build a Semantica knowledge graph from pre-extracted entities. Follow this guide to convert entities to JSON, use GraphTool API, and visualize your graph.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-10

---

**You can build a Semantica knowledge graph by converting pre-extracted entities into JSON-compatible dictionaries, initializing a Model Context Protocol (MCP) session to access the GraphTool API, and using `add_node` and `add_edge` methods to populate the graph before validating with `summary()` and visualizing with `KGVisualizer`.**

The `semantica-agi/semantica` repository provides a programmatic, storage-agnostic API for constructing knowledge graphs from any set of pre-extracted entities—whether they originate from LLM outputs, NER pipelines, or manual curation. This guide walks through the canonical workflow using the actual source implementations.

## Understanding the Three-Phase Construction Workflow

Building a knowledge graph from pre-extracted entities follows three logical phases, each handled by specific components in the Semantica codebase.

### Phase 1: Prepare the Entity Payload

Convert your extracted entities into a list of JSON-compatible dictionaries. Each entity must contain at least an `id`, a human-readable `label`, a `type` (e.g., *Person*, *Organisation*, *Concept*), and optional `metadata` for provenance tracking.

According to the implementation in [`semantica_mcp/mcp/tools/graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/graph.py), the `add_node` method expects these fields to construct graph vertices properly.

### Phase 2: Initialize the Graph and Ingest Data

Open a Semantica MCP session and obtain the GraphTool via `client.tools.graph()`. Use the tool's `add_node` method (defined at line 17 of [`semantica_mcp/mcp/tools/graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/graph.py)) to insert entities and `add_edge` (line 125) to create relationships. For bulk ingestion, the **SeedManager** class (line 101 of [`semantica/seed/seed_manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/seed/seed_manager.py)) can ingest CSV or JSON files directly, handling deduplication and type inference automatically.

### Phase 3: Validate and Visualize the Knowledge Graph

After insertion, retrieve a high-level statistical overview using `graph.summary()` (implemented at line 150 of [`semantica_mcp/mcp/tools/graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/graph.py)). For interactive exploration, instantiate the `KGVisualizer` class from [`semantica/visualization/kg_visualizer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/kg_visualizer.py) (line 18) to render the graph using Plotly.

## Step-by-Step Implementation Guide

### Initialize the MCP Client and Graph Tool

Start by establishing a connection to the Semantica MCP server. The `SemanticaMCP` class in [`semantica_mcp/mcp/session.py`](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/session.py) provides an in-process client for local development.

```python
from semantica_mcp.mcp.session import SemanticaMCP

# Initialize client and obtain the graph tool

client = SemanticaMCP()
graph_tool = client.tools.graph()

```

### Format Your Pre-Extracted Entities

Structure your entities to match the expected schema. This example shows entities that might come from an LLM extraction pipeline or SpaCy NER:

```python
entities = [
    {"id": "e1", "label": "Alice", "type": "Person", "metadata": {"source": "doc1"}},
    {"id": "e2", "label": "Acme Corp", "type": "Company", "metadata": {"source": "doc2"}},
    {"id": "e3", "label": "Quantum AI", "type": "Concept", "metadata": {"source": "doc3"}},
]

```

### Add Entities as Nodes

Iterate through your entity list and call `add_node` for each entry. This method abstracts the underlying storage backend (in-memory, ArangoDB, or Neo4j), keeping your code storage-agnostic.

```python
for ent in entities:
    graph_tool.add_node(
        node_id=ent["id"],
        label=ent["label"],
        node_type=ent["type"],
        metadata=ent["metadata"],
    )

```

### Create Relationships with add_edge

Define relationships between entities using the `add_edge` method. Each edge requires a source ID, target ID, and relationship label:

```python
relationships = [
    {"source": "e1", "target": "e2", "label": "EMPLOYED_BY"},
    {"source": "e2", "target": "e3", "label": "WORKS_ON"},
]

for rel in relationships:
    graph_tool.add_edge(
        source_id=rel["source"],
        target_id=rel["target"],
        label=rel["label"],
    )

```

### Validate Graph Construction

Verify successful ingestion by calling the `summary` method, which returns node counts, edge counts, and type breakdowns:

```python
summary = graph_tool.summary()
print(summary)

```

## Alternative: Bulk Import with SeedManager

For large-scale ingestion, bypass the explicit Python loop and use the **SeedManager** to load flat files directly. This approach handles batch processing, deduplication, and metadata preservation automatically.

```python
from semantica.seed.seed_manager import SeedManager

# Initialize with your MCP client

seed_manager = SeedManager(client)

# Load from CSV (columns: id, label, type, metadata)

seed_manager.load_csv("path/to/entities.csv")

# Or load from JSON

seed_manager.load_json("path/to/entities.json")

```

The `SeedManager` implementation in [`semantica/seed/seed_manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/seed/seed_manager.py) provides efficient bulk ingestion while maintaining data integrity constraints defined in the graph schema.

## Visualizing and Exporting Your Knowledge Graph

### Interactive Visualization

Generate interactive network visualizations using the `KGVisualizer` class. The implementation in [`semantica/visualization/kg_visualizer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/kg_visualizer.py) wraps Plotly and NetworkX for browser-based exploration:

```python
from semantica.visualization.kg_visualizer import KGVisualizer

visualizer = KGVisualizer(client)
visualizer.visualize()  # Opens interactive HTML view

```

For temporal graphs (entities with timestamps), use the `TemporalVisualizer` in [`semantica/visualization/temporal_visualizer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/temporal_visualizer.py) instead.

### Export Utilities

Export your constructed graph to standard formats using the export tools defined in [`semantica_mcp/mcp/tools/export.py`](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/export.py):

- **JSON** for machine-readable interchange
- **CSV** for tabular analysis
- **Parquet** for high-performance analytics
- **GraphML** for compatibility with Gephi and other graph analysis tools

## Summary

- **Prepare entities** as JSON-compatible dictionaries with `id`, `label`, `type`, and optional `metadata` fields
- **Initialize the GraphTool** via `SemanticaMCP().tools.graph()` to access storage-agnostic mutation methods
- **Add nodes** using `add_node()` (line 17 of [`semantica_mcp/mcp/tools/graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/graph.py)) and edges using `add_edge()` (line 125)
- **Bulk load** large datasets using `SeedManager` from [`semantica/seed/seed_manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/seed/seed_manager.py) for automatic deduplication
- **Validate** construction with `summary()` and **visualize** with `KGVisualizer` from [`semantica/visualization/kg_visualizer.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/visualization/kg_visualizer.py)

## Frequently Asked Questions

### What data format do I need for pre-extracted entities?

Pre-extracted entities must be formatted as JSON-compatible dictionaries containing at minimum an `id` (unique identifier), `label` (human-readable name), and `type` (entity category). Optional `metadata` fields can store provenance information, confidence scores, or source document references. The `add_node` method in [`semantica_mcp/mcp/tools/graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/graph.py) accepts these parameters directly.

### Can I use the SeedManager for JSON files as well as CSV?

Yes. The `SeedManager` class in [`semantica/seed/seed_manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/seed/seed_manager.py) provides both `load_csv()` and `load_json()` methods. JSON files should contain an array of objects matching the entity schema, while CSV files should include columns for `id`, `label`, `type`, and `metadata`. The manager automatically handles type inference and entity deduplication during ingestion.

### How do I handle entity deduplication during bulk import?

The `SeedManager` handles deduplication automatically by checking for existing node IDs before insertion. If you are using the low-level `GraphTool` API directly, you should implement your own deduplication logic by querying existing nodes with `search_nodes` before calling `add_node`, or ensure your entity extraction pipeline assigns consistent, deterministic identifiers to avoid duplicates.

### Is the GraphTool storage-agnostic?

Yes. The `GraphTool` implementation in [`semantica_mcp/mcp/tools/graph.py`](https://github.com/semantica-agi/semantica/blob/main/semantica_mcp/mcp/tools/graph.py) abstracts the underlying storage layer, supporting in-memory graphs for development, ArangoDB for production document-graph hybrid storage, and Neo4j for native graph persistence. Your construction code remains identical regardless of the backend configuration, enabling seamless migration from prototyping to production scales.