# How to Construct a Knowledge Graph from Raw Documents Using Semantica's GraphBuilder

> Easily construct a knowledge graph from raw documents with Semantica's GraphBuilder. Learn how this tool handles NER, relationship extraction, and graph persistence.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-09

---

**Semantica's `GraphBuilder` class orchestrates the entire pipeline to construct a knowledge graph from raw documents, handling everything from named entity recognition and relationship extraction to conflict resolution and persistence in graph databases like Neo4j.**

The `GraphBuilder` component in the semantica-agi/semantica repository serves as the primary interface for transforming unstructured text into structured knowledge representations. By chaining specialized extractors, resolvers, and persistence layers, it automates the complex workflow required to build production-ready knowledge graphs from heterogeneous document sources.

## The GraphBuilder Pipeline Architecture

At the core of the library, [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py) implements the `GraphBuilder` class that orchestrates an eight-stage pipeline. According to the Semantica source code, the `build()` method normalizes input structures, coordinates extraction subsystems, and manages the lifecycle of transformation stages from raw text to serialized graph output.

The architecture decouples extraction concerns from storage concerns. While `NERExtractor`, `RelationExtractor`, and `TripletExtractor` handle semantic parsing, the `EntityResolver` and `ConflictDetector` manage data quality, and the optional `GraphStore` interface handles persistence. This modularity allows you to construct knowledge graphs using only the components necessary for your specific use case.

## Step-by-Step Construction Process

When you invoke `builder.build(sources=...)`, the pipeline executes the following stages in sequence:

### Input Normalization and Validation

The `build()` method first normalizes heterogeneous input structures in `semantica/kg/graph_builder.py#L48-L55`. It accepts a single source string, a list of sources, or a dictionary containing pre-extracted `entities` and `relationships`. This flexibility allows the same interface to process raw documents or partially structured data without code changes.

### Entity and Relationship Extraction

If you provide raw text, the pipeline automatically instantiates `NERExtractor`, `RelationExtractor`, and `TripletExtractor` components. As implemented in `semantica/kg/graph_builder.py#L38-L60`, these extractors are cached on first use to avoid expensive repeated model loading. You control the extraction strategy via constructor parameters like `ner_method="ml"` (spaCy-based) or `relation_method="pattern"` (rule-based).

### Entity Resolution and Deduplication

When `merge_entities=True`, the pipeline invokes `EntityResolver` as defined in `semantica/kg/graph_builder.py#L41-L48`. This component deduplicates and merges similar entities using configurable strategies including `fuzzy`, `exact`, or `ml-based` matching. This step prevents duplicate nodes for entities like "Acme Corp" and "Acme Corporation" that refer to the same real-world object.

### Conflict Detection and Resolution

Setting `resolve_conflicts=True` activates the `ConflictDetector` referenced in `semantica/kg/graph_builder.py#L56-L66`. This stage scans the entity list for contradictory information—such as conflicting birth dates for the same person—and attempts automated remediation before the graph is finalized.

### Temporal Metadata Enrichment

The optional `enable_temporal` flag activates temporal graph support as detailed in `semantica/kg/graph_builder.py#L31-L40`. When enabled, the builder tags edges with validity intervals, enabling snapshot creation and point-in-time queries against historical graph states.

### Persistence to Graph Stores

If you provide a `graph_store` instance (such as `Neo4jGraphStore`), the builder persists the constructed graph in two distinct steps: nodes via `add_nodes` followed by edges via `add_edges`, as shown in `semantica/kg/graph_builder.py#L41-L67`. This two-phase commit ensures referential integrity in external databases.

### Output Structure

Regardless of persistence options, the `build()` method returns a standardized dictionary containing three keys: `entities` (resolved node dictionaries), `relationships` (edge dictionaries with optional temporal metadata), and `metadata` (extraction statistics including entity counts, timestamps, and feature flags), as specified in `semantica/kg/graph_builder.py#L27-L34`.

## Implementation Examples

### Ingesting Raw Text Documents

The most common use case involves passing raw text directly to the builder. The `ProgressTracker` and internal logger provide real-time feedback during processing:

```python
from semantica.kg import GraphBuilder

raw_doc = """
Alice founded Acme Corp in 2020. Bob joined Acme Corp as CTO in 2021.
"""

builder = GraphBuilder(
    merge_entities=True,          # Enable deduplication via EntityResolver

    resolve_conflicts=True,      # Activate conflict detection

    enable_temporal=True,        # Tag edges with time intervals

    ner_method="ml",             # Use spaCy-based NERExtractor

    relation_method="pattern",   # Lightweight pattern-based extraction

    triplet_method="pattern",
)

kg = builder.build(sources=raw_doc)

print(f"Entities: {kg['metadata']['num_entities']}")
print(f"Relationships: {kg['metadata']['num_relationships']}")

```

### Processing Pre-Extracted Structures

You can bypass automatic extraction by supplying pre-annotated structures. This is useful when integrating with external NLP pipelines:

```python
from semantica.kg import GraphBuilder

source = {
    "entities": [
        {"id": "alice", "name": "Alice", "type": "PERSON"},
        {"id": "acme", "name": "Acme Corp", "type": "ORG"},
    ],
    "relationships": [
        {"source": "alice", "target": "acme", "type": "FOUNDED", "metadata": {"year": 2020}},
    ],
}

builder = GraphBuilder(merge_entities=False, resolve_conflicts=False)
kg = builder.build(sources=source)

# Visualize using Semantica's visualization layer

from semantica.visualization import KGVisualizer
viz = KGVisualizer()
viz.show(kg)

```

### Integrating with Neo4j

To persist the constructed knowledge graph directly to Neo4j, inject a `GraphStore` implementation during initialization:

```python
from semantica.kg import GraphBuilder
from semantica.graph_store.neo4j import Neo4jGraphStore

store = Neo4jGraphStore(
    uri="bolt://localhost:7687", 
    auth=("neo4j", "password")
)

builder = GraphBuilder(
    graph_store=store, 
    merge_entities=True
)

kg = builder.build(sources=raw_doc)

# Nodes and edges are automatically written to Neo4j

```

## Core Configuration and Customization

The [`semantica/kg/config.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/config.py) file centralizes default configurations for the knowledge graph construction process, including the `unknown_relation_endpoint` setting used when relationship extractors cannot determine edge types. You can override these defaults via constructor arguments to `GraphBuilder` without modifying source files.

For advanced entity resolution, the [`semantica/kg/entity_resolver.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/entity_resolver.py) module implements the fuzzy, exact, and ML-based merging strategies referenced in the pipeline. Similarly, [`semantica/kg/conflicts/conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/conflicts/conflict_detector.py) provides the automated remediation logic used when `resolve_conflicts` is enabled.

## Summary

- **GraphBuilder** in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py) serves as the unified interface for constructing knowledge graphs from raw documents or structured inputs.
- The pipeline automatically caches expensive ML models (NER, relation extraction) to optimize performance across multiple documents.
- **Entity resolution** and **conflict detection** are optional but recommended stages for production data quality, controlled via `merge_entities` and `resolve_conflicts` flags.
- **Temporal support** enables time-aware graph analytics when you set `enable_temporal=True`.
- The output dictionary standardizes access to entities, relationships, and metadata, regardless of whether you persist to external stores like Neo4j.
- All stages are orchestrated within the `build()` method, which handles input normalization and provides real-time progress tracking.

## Frequently Asked Questions

### What input formats does GraphBuilder accept?

`GraphBuilder.build()` accepts three input variants: a single string containing raw text, a list of text strings for batch processing, or a dictionary with pre-extracted `entities` and `relationships` arrays. This design allows you to reuse the same pipeline for both greenfield extraction and integration with existing NLP outputs.

### How does entity merging handle ambiguous matches?

The `EntityResolver` component—accessed when `merge_entities=True`—supports three strategies defined in [`semantica/kg/entity_resolver.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/entity_resolver.py): exact string matching, fuzzy string similarity, and ML-based embedding clustering. You specify the strategy during `GraphBuilder` initialization, and the resolver automatically consolidates matching entities before the conflict detection stage.

### Can I disable automatic extraction and use my own entities?

Yes. Pass a dictionary with `entities` and `relationships` keys directly to the `sources` parameter, and set `ner_method`, `relation_method`, and `triplet_method` to `None` or omit them. The pipeline will skip the `NERExtractor`, `RelationExtractor`, and `TripletExtractor` stages and proceed directly to resolution, conflict detection, and persistence.

### Which graph databases are supported?

The repository includes a concrete `Neo4jGraphStore` implementation in [`semantica/graph_store/neo4j.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/neo4j.py). The `GraphBuilder` accepts any object implementing the `GraphStore` interface via the `graph_store` parameter, allowing extension to RedisGraph or other backends by implementing the `add_nodes` and `add_edges` methods specified in the base interface.