How to Construct a Knowledge Graph from Raw Documents Using Semantica's GraphBuilder

Semantica's GraphBuilder class orchestrates the entire pipeline to construct a knowledge graph from raw documents, handling everything from named entity recognition and relationship extraction to conflict resolution and persistence in graph databases like Neo4j.

The GraphBuilder component in the semantica-agi/semantica repository serves as the primary interface for transforming unstructured text into structured knowledge representations. By chaining specialized extractors, resolvers, and persistence layers, it automates the complex workflow required to build production-ready knowledge graphs from heterogeneous document sources.

The GraphBuilder Pipeline Architecture

At the core of the library, semantica/kg/graph_builder.py implements the GraphBuilder class that orchestrates an eight-stage pipeline. According to the Semantica source code, the build() method normalizes input structures, coordinates extraction subsystems, and manages the lifecycle of transformation stages from raw text to serialized graph output.

The architecture decouples extraction concerns from storage concerns. While NERExtractor, RelationExtractor, and TripletExtractor handle semantic parsing, the EntityResolver and ConflictDetector manage data quality, and the optional GraphStore interface handles persistence. This modularity allows you to construct knowledge graphs using only the components necessary for your specific use case.

Step-by-Step Construction Process

When you invoke builder.build(sources=...), the pipeline executes the following stages in sequence:

Input Normalization and Validation

The build() method first normalizes heterogeneous input structures in semantica/kg/graph_builder.py#L48-L55. It accepts a single source string, a list of sources, or a dictionary containing pre-extracted entities and relationships. This flexibility allows the same interface to process raw documents or partially structured data without code changes.

Entity and Relationship Extraction

If you provide raw text, the pipeline automatically instantiates NERExtractor, RelationExtractor, and TripletExtractor components. As implemented in semantica/kg/graph_builder.py#L38-L60, these extractors are cached on first use to avoid expensive repeated model loading. You control the extraction strategy via constructor parameters like ner_method="ml" (spaCy-based) or relation_method="pattern" (rule-based).

Entity Resolution and Deduplication

When merge_entities=True, the pipeline invokes EntityResolver as defined in semantica/kg/graph_builder.py#L41-L48. This component deduplicates and merges similar entities using configurable strategies including fuzzy, exact, or ml-based matching. This step prevents duplicate nodes for entities like "Acme Corp" and "Acme Corporation" that refer to the same real-world object.

Conflict Detection and Resolution

Setting resolve_conflicts=True activates the ConflictDetector referenced in semantica/kg/graph_builder.py#L56-L66. This stage scans the entity list for contradictory information—such as conflicting birth dates for the same person—and attempts automated remediation before the graph is finalized.

Temporal Metadata Enrichment

The optional enable_temporal flag activates temporal graph support as detailed in semantica/kg/graph_builder.py#L31-L40. When enabled, the builder tags edges with validity intervals, enabling snapshot creation and point-in-time queries against historical graph states.

Persistence to Graph Stores

If you provide a graph_store instance (such as Neo4jGraphStore), the builder persists the constructed graph in two distinct steps: nodes via add_nodes followed by edges via add_edges, as shown in semantica/kg/graph_builder.py#L41-L67. This two-phase commit ensures referential integrity in external databases.

Output Structure

Regardless of persistence options, the build() method returns a standardized dictionary containing three keys: entities (resolved node dictionaries), relationships (edge dictionaries with optional temporal metadata), and metadata (extraction statistics including entity counts, timestamps, and feature flags), as specified in semantica/kg/graph_builder.py#L27-L34.

Implementation Examples

Ingesting Raw Text Documents

The most common use case involves passing raw text directly to the builder. The ProgressTracker and internal logger provide real-time feedback during processing:

from semantica.kg import GraphBuilder

raw_doc = """
Alice founded Acme Corp in 2020. Bob joined Acme Corp as CTO in 2021.
"""

builder = GraphBuilder(
    merge_entities=True,          # Enable deduplication via EntityResolver

    resolve_conflicts=True,      # Activate conflict detection

    enable_temporal=True,        # Tag edges with time intervals

    ner_method="ml",             # Use spaCy-based NERExtractor

    relation_method="pattern",   # Lightweight pattern-based extraction

    triplet_method="pattern",
)

kg = builder.build(sources=raw_doc)

print(f"Entities: {kg['metadata']['num_entities']}")
print(f"Relationships: {kg['metadata']['num_relationships']}")

Processing Pre-Extracted Structures

You can bypass automatic extraction by supplying pre-annotated structures. This is useful when integrating with external NLP pipelines:

from semantica.kg import GraphBuilder

source = {
    "entities": [
        {"id": "alice", "name": "Alice", "type": "PERSON"},
        {"id": "acme", "name": "Acme Corp", "type": "ORG"},
    ],
    "relationships": [
        {"source": "alice", "target": "acme", "type": "FOUNDED", "metadata": {"year": 2020}},
    ],
}

builder = GraphBuilder(merge_entities=False, resolve_conflicts=False)
kg = builder.build(sources=source)

# Visualize using Semantica's visualization layer

from semantica.visualization import KGVisualizer
viz = KGVisualizer()
viz.show(kg)

Integrating with Neo4j

To persist the constructed knowledge graph directly to Neo4j, inject a GraphStore implementation during initialization:

from semantica.kg import GraphBuilder
from semantica.graph_store.neo4j import Neo4jGraphStore

store = Neo4jGraphStore(
    uri="bolt://localhost:7687", 
    auth=("neo4j", "password")
)

builder = GraphBuilder(
    graph_store=store, 
    merge_entities=True
)

kg = builder.build(sources=raw_doc)

# Nodes and edges are automatically written to Neo4j

Core Configuration and Customization

The semantica/kg/config.py file centralizes default configurations for the knowledge graph construction process, including the unknown_relation_endpoint setting used when relationship extractors cannot determine edge types. You can override these defaults via constructor arguments to GraphBuilder without modifying source files.

For advanced entity resolution, the semantica/kg/entity_resolver.py module implements the fuzzy, exact, and ML-based merging strategies referenced in the pipeline. Similarly, semantica/kg/conflicts/conflict_detector.py provides the automated remediation logic used when resolve_conflicts is enabled.

Summary

  • GraphBuilder in semantica/kg/graph_builder.py serves as the unified interface for constructing knowledge graphs from raw documents or structured inputs.
  • The pipeline automatically caches expensive ML models (NER, relation extraction) to optimize performance across multiple documents.
  • Entity resolution and conflict detection are optional but recommended stages for production data quality, controlled via merge_entities and resolve_conflicts flags.
  • Temporal support enables time-aware graph analytics when you set enable_temporal=True.
  • The output dictionary standardizes access to entities, relationships, and metadata, regardless of whether you persist to external stores like Neo4j.
  • All stages are orchestrated within the build() method, which handles input normalization and provides real-time progress tracking.

Frequently Asked Questions

What input formats does GraphBuilder accept?

GraphBuilder.build() accepts three input variants: a single string containing raw text, a list of text strings for batch processing, or a dictionary with pre-extracted entities and relationships arrays. This design allows you to reuse the same pipeline for both greenfield extraction and integration with existing NLP outputs.

How does entity merging handle ambiguous matches?

The EntityResolver component—accessed when merge_entities=True—supports three strategies defined in semantica/kg/entity_resolver.py: exact string matching, fuzzy string similarity, and ML-based embedding clustering. You specify the strategy during GraphBuilder initialization, and the resolver automatically consolidates matching entities before the conflict detection stage.

Can I disable automatic extraction and use my own entities?

Yes. Pass a dictionary with entities and relationships keys directly to the sources parameter, and set ner_method, relation_method, and triplet_method to None or omit them. The pipeline will skip the NERExtractor, RelationExtractor, and TripletExtractor stages and proceed directly to resolution, conflict detection, and persistence.

Which graph databases are supported?

The repository includes a concrete Neo4jGraphStore implementation in semantica/graph_store/neo4j.py. The GraphBuilder accepts any object implementing the GraphStore interface via the graph_store parameter, allowing extension to RedisGraph or other backends by implementing the add_nodes and add_edges methods specified in the base interface.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →