How to Build a Knowledge Graph with Semantica from Raw Documents: A Complete Pipeline Guide
Semantica transforms unstructured text into a structured knowledge graph by orchestrating ingestion, parsing, semantic extraction, and graph assembly through a fault-tolerant pipeline DSL that executes as a directed acyclic graph (DAG).
Building a knowledge graph from raw documents requires coordinating multiple data transformations, from entity recognition to relationship extraction and deduplication. Semantica provides a modular Python framework that wires these stages together using a declarative pipeline architecture. In this guide, you will learn how to build a knowledge graph with Semantica using both high-level templates and low-level pipeline APIs according to the semantica-agi/semantica source code.
Understanding the Knowledge Graph Pipeline Architecture
Semantica’s knowledge graph construction follows a deterministic seven-stage flow. Each stage is handled by a specific component class that can be customized or replaced.
-
Ingestion – Load raw documents using an Ingestor (e.g.,
FileIngestor) from files, URLs, or database rows. -
Parsing – Convert binary or text inputs into structured document objects via
DocumentParser. -
Normalization – Clean and standardize text using
TextNormalizerto ensure consistent downstream processing. -
Semantic Extraction – Identify entities, relations, and attributes using the semantic-extract module. The
NERExtractorandRelationExtractorclasses form the foundation of the knowledge graph schema. -
Deduplication & Entity Resolution – Collapse duplicate mentions and resolve ambiguous references using
EntityResolverand deduplication algorithms implemented insemantica/kg/entity_resolver.py. -
Graph Building – Assemble cleaned triples into a graph structure using
GraphBuilder. This component supports temporal tracking, conflict resolution strategies, and custom handling for unknown endpoints. -
Optional Analytics – Execute graph algorithms including node embeddings, similarity calculations, path-finding, and community detection using utilities in
semantica/kg/algorithms/.
The pipeline executes via the ExecutionEngine class, which provides parallel workers, retry policies, and progress tracking as defined in semantica/pipeline.py.
Method 1: Quick Start with the kg_construction Template
The fastest way to build a knowledge graph uses the PipelineTemplateManager located in semantica/pipeline/template_manager.py. This manager provides a pre-configured kg_construction template that wires all seven stages in the correct order.
from semantica.pipeline import PipelineTemplateManager, ExecutionEngine
from semantica.ingest import FileIngestor
from semantica.parse import DocumentParser
from semantica.semantic_extract import NERExtractor, RelationExtractor
from semantica.kg import GraphBuilder
# Initialize the template manager
manager = PipelineTemplateManager()
builder = manager.create_pipeline_from_template("kg_construction")
# Configure custom handlers (optional)
ingestor = FileIngestor()
parser = DocumentParser()
ner_ext = NERExtractor(method="ml")
rel_ext = RelationExtractor(method="ml")
kg_builder = GraphBuilder(merge_entities=True, resolve_conflicts=True)
# Override default handlers with specific implementations
builder.set_handler("ingest", ingestor.ingest_file)
builder.set_handler("parse", parser.parse)
builder.set_handler("extract", ner_ext.extract)
builder.set_handler("rel_extract", rel_ext.extract)
builder.set_handler("build_kg", kg_builder.build)
# Execute with parallel processing
pipeline = builder.build("my_kg_pipeline")
engine = ExecutionEngine(max_workers=4)
result = engine.execute_pipeline(pipeline, data="path/to/docs/")
kg = result.output
print(f"KG contains {len(kg['entities'])} entities and {len(kg['relationships'])} relationships")
The template automatically sequences operations as ingest → parse → normalize → extract → rel_extract → build_kg → deduplicate → export. You only need to override handlers where custom logic is required.
Method 2: Manual Pipeline Construction for Custom Workflows
For workflows requiring non-linear processing or conditional branching, use the PipelineBuilder class directly. This approach defines the DAG structure explicitly via connect_steps().
from semantica.pipeline import PipelineBuilder, ExecutionEngine
from semantica.ingest import FileIngestor
from semantica.parse import DocumentParser
from semantica.semantic_extract import NERExtractor, RelationExtractor
from semantica.kg import GraphBuilder
# Instantiate handlers
ingestor = FileIngestor()
parser = DocumentParser()
ner_ext = NERExtractor(method="ml")
rel_ext = RelationExtractor(method="ml")
kg_builder = GraphBuilder(
merge_entities=True,
entity_resolution_strategy="fuzzy",
resolve_conflicts=True
)
# Construct the DAG manually
builder = PipelineBuilder()
builder.add_step("ingest", "file_ingest", handler=ingestor.ingest_file)
builder.add_step("parse", "document_parse", handler=parser.parse)
builder.add_step("extract", "ner_extract", handler=ner_ext.extract)
builder.add_step("rel_ext", "rel_extract", handler=rel_ext.extract)
builder.add_step("build_kg", "graph_build", handler=kg_builder.build)
# Define execution flow
builder.connect_steps("ingest", "parse")
builder.connect_steps("parse", "extract")
builder.connect_steps("extract", "rel_ext")
builder.connect_steps("rel_ext", "build_kg")
# Build and execute
pipeline = builder.build("manual_kg_pipeline")
engine = ExecutionEngine()
result = engine.execute_pipeline(pipeline, data="data/")
kg = result.output
print(f"KG built with {len(kg['entities'])} entities.")
Manual construction provides granular control over step dependencies and allows insertion of custom processing nodes between standard stages.
Adding Graph Analytics and Embeddings
Once the knowledge graph is constructed, apply graph algorithms using classes from semantica/kg/algorithms/. The following example computes node embeddings and similarity scores.
from semantica.kg import NodeEmbedder
# kg is the output from previous pipeline execution
embedder = NodeEmbedder(method="node2vec", embedding_dimension=128)
embeddings = embedder.compute_embeddings(
graph_store=kg,
node_labels=["Entity", "Person"],
relationship_types=["RELATED_TO", "KNOWS"]
)
# Retrieve top-k similar nodes
similar = embedder.find_similar_nodes(
graph_store=kg,
target_node="entity_42",
top_k=5
)
print(f"Nodes most similar to entity_42:")
for node_id, score in similar:
print(f" {node_id}: {score:.4f}")
The NodeEmbedder supports multiple algorithms including node2vec and can filter by specific node labels and relationship types defined in your graph schema.
Key Configuration Options for GraphBuilder
The GraphBuilder class in semantica/kg/graph_builder.py accepts several parameters that control knowledge graph construction behavior:
merge_entities(bool): Enables entity consolidation across documents when set toTrue.resolve_conflicts(bool): Activates conflict resolution for contradictory relationship assertions.entity_resolution_strategy(str): Specifies the matching algorithm for deduplication, such as"fuzzy"or"exact".temporal_tracking(bool): Records timestamps for entity and relationship creation when enabled.
These options are passed during instantiation and affect how the builder processes triples during the build() method execution.
Summary
- Semantica builds knowledge graphs through a seven-stage pipeline: ingestion, parsing, normalization, semantic extraction, deduplication, graph building, and optional analytics.
- Use
PipelineTemplateManagerwith the"kg_construction"template for standard workflows, referencingsemantica/pipeline/template_manager.py. - For custom flows, manually construct DAGs using
PipelineBuilderandconnect_steps()as implemented insemantica/pipeline.py. - Configure entity resolution and conflict handling via
GraphBuilderparameters likeentity_resolution_strategyandmerge_entities. - Execute pipelines with
ExecutionEngineto leverage parallel workers and fault-tolerant retry policies. - Extend functionality with graph algorithms from
semantica/kg/algorithms/including embedding generators and similarity calculators.
Frequently Asked Questions
What file formats does Semantica support for document ingestion?
The FileIngestor class handles multiple formats including plain text, PDF, and structured data files. You can extend support by implementing custom ingestors that conform to the base Ingestor interface defined in the ingestion module. The extracted content is passed as a standardized document object to the DocumentParser.
How does Semantica handle duplicate entities across multiple documents?
Duplicate detection occurs in the deduplication stage using EntityResolver from semantica/kg/entity_resolver.py. When GraphBuilder is configured with merge_entities=True, the system collapses mentions referring to the same real-world entity using the strategy specified by entity_resolution_strategy, such as fuzzy matching on entity names or attributes.
Can I customize the machine learning models used for entity and relationship extraction?
Yes. Both NERExtractor and RelationExtractor accept a method parameter during instantiation. Set method="ml" to use machine learning models, or specify alternative approaches supported by the semantic_extract module. You can also inject custom extraction handlers by subclassing the base extractor classes defined in semantica/semantic_extract/__init__.py.
What is the difference between using the template manager versus manual pipeline construction?
The PipelineTemplateManager provides a pre-validated DAG configuration optimized for standard knowledge graph construction, reducing boilerplate code. Manual construction with PipelineBuilder exposes the full DAG API, allowing custom step ordering, conditional branching, and insertion of non-standard processing nodes. Both approaches use the same ExecutionEngine for fault-tolerant, parallel execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →