How to Build a Knowledge Graph with Semantica from Raw Documents: A Complete Pipeline Guide

Semantica transforms unstructured text into a structured knowledge graph by orchestrating ingestion, parsing, semantic extraction, and graph assembly through a fault-tolerant pipeline DSL that executes as a directed acyclic graph (DAG).

Building a knowledge graph from raw documents requires coordinating multiple data transformations, from entity recognition to relationship extraction and deduplication. Semantica provides a modular Python framework that wires these stages together using a declarative pipeline architecture. In this guide, you will learn how to build a knowledge graph with Semantica using both high-level templates and low-level pipeline APIs according to the semantica-agi/semantica source code.

Understanding the Knowledge Graph Pipeline Architecture

Semantica’s knowledge graph construction follows a deterministic seven-stage flow. Each stage is handled by a specific component class that can be customized or replaced.

  1. Ingestion – Load raw documents using an Ingestor (e.g., FileIngestor) from files, URLs, or database rows.

  2. Parsing – Convert binary or text inputs into structured document objects via DocumentParser.

  3. Normalization – Clean and standardize text using TextNormalizer to ensure consistent downstream processing.

  4. Semantic Extraction – Identify entities, relations, and attributes using the semantic-extract module. The NERExtractor and RelationExtractor classes form the foundation of the knowledge graph schema.

  5. Deduplication & Entity Resolution – Collapse duplicate mentions and resolve ambiguous references using EntityResolver and deduplication algorithms implemented in semantica/kg/entity_resolver.py.

  6. Graph Building – Assemble cleaned triples into a graph structure using GraphBuilder. This component supports temporal tracking, conflict resolution strategies, and custom handling for unknown endpoints.

  7. Optional Analytics – Execute graph algorithms including node embeddings, similarity calculations, path-finding, and community detection using utilities in semantica/kg/algorithms/.

The pipeline executes via the ExecutionEngine class, which provides parallel workers, retry policies, and progress tracking as defined in semantica/pipeline.py.

Method 1: Quick Start with the kg_construction Template

The fastest way to build a knowledge graph uses the PipelineTemplateManager located in semantica/pipeline/template_manager.py. This manager provides a pre-configured kg_construction template that wires all seven stages in the correct order.

from semantica.pipeline import PipelineTemplateManager, ExecutionEngine
from semantica.ingest import FileIngestor
from semantica.parse import DocumentParser
from semantica.semantic_extract import NERExtractor, RelationExtractor
from semantica.kg import GraphBuilder

# Initialize the template manager

manager = PipelineTemplateManager()
builder = manager.create_pipeline_from_template("kg_construction")

# Configure custom handlers (optional)

ingestor = FileIngestor()
parser = DocumentParser()
ner_ext = NERExtractor(method="ml")
rel_ext = RelationExtractor(method="ml")
kg_builder = GraphBuilder(merge_entities=True, resolve_conflicts=True)

# Override default handlers with specific implementations

builder.set_handler("ingest", ingestor.ingest_file)
builder.set_handler("parse", parser.parse)
builder.set_handler("extract", ner_ext.extract)
builder.set_handler("rel_extract", rel_ext.extract)
builder.set_handler("build_kg", kg_builder.build)

# Execute with parallel processing

pipeline = builder.build("my_kg_pipeline")
engine = ExecutionEngine(max_workers=4)
result = engine.execute_pipeline(pipeline, data="path/to/docs/")

kg = result.output
print(f"KG contains {len(kg['entities'])} entities and {len(kg['relationships'])} relationships")

The template automatically sequences operations as ingest → parse → normalize → extract → rel_extract → build_kg → deduplicate → export. You only need to override handlers where custom logic is required.

Method 2: Manual Pipeline Construction for Custom Workflows

For workflows requiring non-linear processing or conditional branching, use the PipelineBuilder class directly. This approach defines the DAG structure explicitly via connect_steps().

from semantica.pipeline import PipelineBuilder, ExecutionEngine
from semantica.ingest import FileIngestor
from semantica.parse import DocumentParser
from semantica.semantic_extract import NERExtractor, RelationExtractor
from semantica.kg import GraphBuilder

# Instantiate handlers

ingestor = FileIngestor()
parser = DocumentParser()
ner_ext = NERExtractor(method="ml")
rel_ext = RelationExtractor(method="ml")
kg_builder = GraphBuilder(
    merge_entities=True,
    entity_resolution_strategy="fuzzy",
    resolve_conflicts=True
)

# Construct the DAG manually

builder = PipelineBuilder()
builder.add_step("ingest", "file_ingest", handler=ingestor.ingest_file)
builder.add_step("parse", "document_parse", handler=parser.parse)
builder.add_step("extract", "ner_extract", handler=ner_ext.extract)
builder.add_step("rel_ext", "rel_extract", handler=rel_ext.extract)
builder.add_step("build_kg", "graph_build", handler=kg_builder.build)

# Define execution flow

builder.connect_steps("ingest", "parse")
builder.connect_steps("parse", "extract")
builder.connect_steps("extract", "rel_ext")
builder.connect_steps("rel_ext", "build_kg")

# Build and execute

pipeline = builder.build("manual_kg_pipeline")
engine = ExecutionEngine()
result = engine.execute_pipeline(pipeline, data="data/")

kg = result.output
print(f"KG built with {len(kg['entities'])} entities.")

Manual construction provides granular control over step dependencies and allows insertion of custom processing nodes between standard stages.

Adding Graph Analytics and Embeddings

Once the knowledge graph is constructed, apply graph algorithms using classes from semantica/kg/algorithms/. The following example computes node embeddings and similarity scores.

from semantica.kg import NodeEmbedder

# kg is the output from previous pipeline execution

embedder = NodeEmbedder(method="node2vec", embedding_dimension=128)
embeddings = embedder.compute_embeddings(
    graph_store=kg,
    node_labels=["Entity", "Person"],
    relationship_types=["RELATED_TO", "KNOWS"]
)

# Retrieve top-k similar nodes

similar = embedder.find_similar_nodes(
    graph_store=kg,
    target_node="entity_42",
    top_k=5
)

print(f"Nodes most similar to entity_42:")
for node_id, score in similar:
    print(f"  {node_id}: {score:.4f}")

The NodeEmbedder supports multiple algorithms including node2vec and can filter by specific node labels and relationship types defined in your graph schema.

Key Configuration Options for GraphBuilder

The GraphBuilder class in semantica/kg/graph_builder.py accepts several parameters that control knowledge graph construction behavior:

  • merge_entities (bool): Enables entity consolidation across documents when set to True.
  • resolve_conflicts (bool): Activates conflict resolution for contradictory relationship assertions.
  • entity_resolution_strategy (str): Specifies the matching algorithm for deduplication, such as "fuzzy" or "exact".
  • temporal_tracking (bool): Records timestamps for entity and relationship creation when enabled.

These options are passed during instantiation and affect how the builder processes triples during the build() method execution.

Summary

  • Semantica builds knowledge graphs through a seven-stage pipeline: ingestion, parsing, normalization, semantic extraction, deduplication, graph building, and optional analytics.
  • Use PipelineTemplateManager with the "kg_construction" template for standard workflows, referencing semantica/pipeline/template_manager.py.
  • For custom flows, manually construct DAGs using PipelineBuilder and connect_steps() as implemented in semantica/pipeline.py.
  • Configure entity resolution and conflict handling via GraphBuilder parameters like entity_resolution_strategy and merge_entities.
  • Execute pipelines with ExecutionEngine to leverage parallel workers and fault-tolerant retry policies.
  • Extend functionality with graph algorithms from semantica/kg/algorithms/ including embedding generators and similarity calculators.

Frequently Asked Questions

What file formats does Semantica support for document ingestion?

The FileIngestor class handles multiple formats including plain text, PDF, and structured data files. You can extend support by implementing custom ingestors that conform to the base Ingestor interface defined in the ingestion module. The extracted content is passed as a standardized document object to the DocumentParser.

How does Semantica handle duplicate entities across multiple documents?

Duplicate detection occurs in the deduplication stage using EntityResolver from semantica/kg/entity_resolver.py. When GraphBuilder is configured with merge_entities=True, the system collapses mentions referring to the same real-world entity using the strategy specified by entity_resolution_strategy, such as fuzzy matching on entity names or attributes.

Can I customize the machine learning models used for entity and relationship extraction?

Yes. Both NERExtractor and RelationExtractor accept a method parameter during instantiation. Set method="ml" to use machine learning models, or specify alternative approaches supported by the semantic_extract module. You can also inject custom extraction handlers by subclassing the base extractor classes defined in semantica/semantic_extract/__init__.py.

What is the difference between using the template manager versus manual pipeline construction?

The PipelineTemplateManager provides a pre-validated DAG configuration optimized for standard knowledge graph construction, reducing boilerplate code. Manual construction with PipelineBuilder exposes the full DAG API, allowing custom step ordering, conditional branching, and insertion of non-standard processing nodes. Both approaches use the same ExecutionEngine for fault-tolerant, parallel execution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →