# How to Build a Knowledge Graph with Semantica from Raw Documents: A Complete Pipeline Guide

> Learn how to build a knowledge graph from raw documents using Semantica. Discover a complete pipeline guide for transforming unstructured text into structured data.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-10

---

**Semantica transforms unstructured text into a structured knowledge graph by orchestrating ingestion, parsing, semantic extraction, and graph assembly through a fault-tolerant pipeline DSL that executes as a directed acyclic graph (DAG).**

Building a knowledge graph from raw documents requires coordinating multiple data transformations, from entity recognition to relationship extraction and deduplication. Semantica provides a modular Python framework that wires these stages together using a declarative pipeline architecture. In this guide, you will learn how to build a knowledge graph with Semantica using both high-level templates and low-level pipeline APIs according to the `semantica-agi/semantica` source code.

## Understanding the Knowledge Graph Pipeline Architecture

Semantica’s knowledge graph construction follows a deterministic seven-stage flow. Each stage is handled by a specific component class that can be customized or replaced.

1. **Ingestion** – Load raw documents using an **Ingestor** (e.g., `FileIngestor`) from files, URLs, or database rows.

2. **Parsing** – Convert binary or text inputs into structured document objects via `DocumentParser`.

3. **Normalization** – Clean and standardize text using `TextNormalizer` to ensure consistent downstream processing.

4. **Semantic Extraction** – Identify entities, relations, and attributes using the **semantic-extract** module. The `NERExtractor` and `RelationExtractor` classes form the foundation of the knowledge graph schema.

5. **Deduplication & Entity Resolution** – Collapse duplicate mentions and resolve ambiguous references using `EntityResolver` and deduplication algorithms implemented in [`semantica/kg/entity_resolver.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/entity_resolver.py).

6. **Graph Building** – Assemble cleaned triples into a graph structure using `GraphBuilder`. This component supports temporal tracking, conflict resolution strategies, and custom handling for unknown endpoints.

7. **Optional Analytics** – Execute graph algorithms including node embeddings, similarity calculations, path-finding, and community detection using utilities in `semantica/kg/algorithms/`.

The pipeline executes via the `ExecutionEngine` class, which provides parallel workers, retry policies, and progress tracking as defined in [`semantica/pipeline.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline.py).

## Method 1: Quick Start with the kg_construction Template

The fastest way to build a knowledge graph uses the `PipelineTemplateManager` located in [`semantica/pipeline/template_manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/template_manager.py). This manager provides a pre-configured `kg_construction` template that wires all seven stages in the correct order.

```python
from semantica.pipeline import PipelineTemplateManager, ExecutionEngine
from semantica.ingest import FileIngestor
from semantica.parse import DocumentParser
from semantica.semantic_extract import NERExtractor, RelationExtractor
from semantica.kg import GraphBuilder

# Initialize the template manager

manager = PipelineTemplateManager()
builder = manager.create_pipeline_from_template("kg_construction")

# Configure custom handlers (optional)

ingestor = FileIngestor()
parser = DocumentParser()
ner_ext = NERExtractor(method="ml")
rel_ext = RelationExtractor(method="ml")
kg_builder = GraphBuilder(merge_entities=True, resolve_conflicts=True)

# Override default handlers with specific implementations

builder.set_handler("ingest", ingestor.ingest_file)
builder.set_handler("parse", parser.parse)
builder.set_handler("extract", ner_ext.extract)
builder.set_handler("rel_extract", rel_ext.extract)
builder.set_handler("build_kg", kg_builder.build)

# Execute with parallel processing

pipeline = builder.build("my_kg_pipeline")
engine = ExecutionEngine(max_workers=4)
result = engine.execute_pipeline(pipeline, data="path/to/docs/")

kg = result.output
print(f"KG contains {len(kg['entities'])} entities and {len(kg['relationships'])} relationships")

```

The template automatically sequences operations as `ingest → parse → normalize → extract → rel_extract → build_kg → deduplicate → export`. You only need to override handlers where custom logic is required.

## Method 2: Manual Pipeline Construction for Custom Workflows

For workflows requiring non-linear processing or conditional branching, use the `PipelineBuilder` class directly. This approach defines the DAG structure explicitly via `connect_steps()`.

```python
from semantica.pipeline import PipelineBuilder, ExecutionEngine
from semantica.ingest import FileIngestor
from semantica.parse import DocumentParser
from semantica.semantic_extract import NERExtractor, RelationExtractor
from semantica.kg import GraphBuilder

# Instantiate handlers

ingestor = FileIngestor()
parser = DocumentParser()
ner_ext = NERExtractor(method="ml")
rel_ext = RelationExtractor(method="ml")
kg_builder = GraphBuilder(
    merge_entities=True,
    entity_resolution_strategy="fuzzy",
    resolve_conflicts=True
)

# Construct the DAG manually

builder = PipelineBuilder()
builder.add_step("ingest", "file_ingest", handler=ingestor.ingest_file)
builder.add_step("parse", "document_parse", handler=parser.parse)
builder.add_step("extract", "ner_extract", handler=ner_ext.extract)
builder.add_step("rel_ext", "rel_extract", handler=rel_ext.extract)
builder.add_step("build_kg", "graph_build", handler=kg_builder.build)

# Define execution flow

builder.connect_steps("ingest", "parse")
builder.connect_steps("parse", "extract")
builder.connect_steps("extract", "rel_ext")
builder.connect_steps("rel_ext", "build_kg")

# Build and execute

pipeline = builder.build("manual_kg_pipeline")
engine = ExecutionEngine()
result = engine.execute_pipeline(pipeline, data="data/")

kg = result.output
print(f"KG built with {len(kg['entities'])} entities.")

```

Manual construction provides granular control over step dependencies and allows insertion of custom processing nodes between standard stages.

## Adding Graph Analytics and Embeddings

Once the knowledge graph is constructed, apply graph algorithms using classes from `semantica/kg/algorithms/`. The following example computes node embeddings and similarity scores.

```python
from semantica.kg import NodeEmbedder

# kg is the output from previous pipeline execution

embedder = NodeEmbedder(method="node2vec", embedding_dimension=128)
embeddings = embedder.compute_embeddings(
    graph_store=kg,
    node_labels=["Entity", "Person"],
    relationship_types=["RELATED_TO", "KNOWS"]
)

# Retrieve top-k similar nodes

similar = embedder.find_similar_nodes(
    graph_store=kg,
    target_node="entity_42",
    top_k=5
)

print(f"Nodes most similar to entity_42:")
for node_id, score in similar:
    print(f"  {node_id}: {score:.4f}")

```

The `NodeEmbedder` supports multiple algorithms including node2vec and can filter by specific node labels and relationship types defined in your graph schema.

## Key Configuration Options for GraphBuilder

The `GraphBuilder` class in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py) accepts several parameters that control knowledge graph construction behavior:

- **`merge_entities`** (`bool`): Enables entity consolidation across documents when set to `True`.
- **`resolve_conflicts`** (`bool`): Activates conflict resolution for contradictory relationship assertions.
- **`entity_resolution_strategy`** (`str`): Specifies the matching algorithm for deduplication, such as `"fuzzy"` or `"exact"`.
- **`temporal_tracking`** (`bool`): Records timestamps for entity and relationship creation when enabled.

These options are passed during instantiation and affect how the builder processes triples during the `build()` method execution.

## Summary

- **Semantica** builds knowledge graphs through a seven-stage pipeline: ingestion, parsing, normalization, semantic extraction, deduplication, graph building, and optional analytics.
- Use `PipelineTemplateManager` with the `"kg_construction"` template for standard workflows, referencing [`semantica/pipeline/template_manager.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline/template_manager.py).
- For custom flows, manually construct DAGs using `PipelineBuilder` and `connect_steps()` as implemented in [`semantica/pipeline.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/pipeline.py).
- Configure entity resolution and conflict handling via `GraphBuilder` parameters like `entity_resolution_strategy` and `merge_entities`.
- Execute pipelines with `ExecutionEngine` to leverage parallel workers and fault-tolerant retry policies.
- Extend functionality with graph algorithms from `semantica/kg/algorithms/` including embedding generators and similarity calculators.

## Frequently Asked Questions

### What file formats does Semantica support for document ingestion?

The `FileIngestor` class handles multiple formats including plain text, PDF, and structured data files. You can extend support by implementing custom ingestors that conform to the base `Ingestor` interface defined in the ingestion module. The extracted content is passed as a standardized document object to the `DocumentParser`.

### How does Semantica handle duplicate entities across multiple documents?

Duplicate detection occurs in the deduplication stage using `EntityResolver` from [`semantica/kg/entity_resolver.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/entity_resolver.py). When `GraphBuilder` is configured with `merge_entities=True`, the system collapses mentions referring to the same real-world entity using the strategy specified by `entity_resolution_strategy`, such as fuzzy matching on entity names or attributes.

### Can I customize the machine learning models used for entity and relationship extraction?

Yes. Both `NERExtractor` and `RelationExtractor` accept a `method` parameter during instantiation. Set `method="ml"` to use machine learning models, or specify alternative approaches supported by the `semantic_extract` module. You can also inject custom extraction handlers by subclassing the base extractor classes defined in [`semantica/semantic_extract/__init__.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/semantic_extract/__init__.py).

### What is the difference between using the template manager versus manual pipeline construction?

The `PipelineTemplateManager` provides a pre-validated DAG configuration optimized for standard knowledge graph construction, reducing boilerplate code. Manual construction with `PipelineBuilder` exposes the full DAG API, allowing custom step ordering, conditional branching, and insertion of non-standard processing nodes. Both approaches use the same `ExecutionEngine` for fault-tolerant, parallel execution.