# How Semantica Transforms Unstructured Data into a Knowledge Graph

> Discover how Semantica transforms unstructured data into a knowledge graph using a multi-stage pipeline. Learn about entity extraction, normalization, and graph assembly.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-08

---

**Semantica converts unstructured data into structured knowledge graphs through a multi-stage pipeline that includes ingestion, entity extraction, normalization, deduplication, and graph assembly, implemented primarily in the `GraphBuilder` class.**

The open-source Semantica framework (semantica-agi/semantica) provides a robust engine to transform unstructured data into a knowledge graph without requiring external APIs by default. By orchestrating specialized ingestors, local extractors, and resolution algorithms, Semantica converts raw text, files, and web sources into query-ready graph structures.

## Unified Ingestion Across Heterogeneous Sources

The pipeline begins with the public `ingest` function, which auto-detects source types and routes data to the appropriate *Ingestor* class. According to [`semantica/ingest/ingest_usage.md`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/ingest_usage.md) (lines 30-45), the system supports `FileIngestor`, `WebIngestor`, `PublicAPIIngestor`, and other specialized handlers to process files, web pages, feeds, streams, and databases.

```python
from semantica.ingest import ingest

# Auto-detect and ingest a PDF file

result = ingest("reports/annual_report.pdf", source_type="file")

```

## Multi-Strategy Text Extraction

For textual sources, `GraphBuilder._extract_from_text` in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py) (lines 48-66) orchestrates three local extractors:

- **NERExtractor** identifies named entities
- **RelationExtractor** extracts explicit relationships (optional)
- **TripletExtractor** extracts subject-predicate-object triplets

All three support configurable backends—`ml`, `pattern`, or `llm`—and operate without external API dependencies by default.

```python
from semantica.kg import GraphBuilder

builder = GraphBuilder(merge_entities=True, resolve_conflicts=True)
graph = builder.build(
    "The Acme Corp acquired Beta Ltd in 2023. Acme's CEO is Jane Doe.",
    extract=True,
    ner_method="ml",
    extract_relations=False
)

```

## Normalization and Entity Resolution

The `_process_item` helper (lines 68-118 in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py)) normalizes inputs—whether strings, Entity objects, Relation objects, or raw dictionaries—into canonical entity or relationship dictionaries. It also promotes synthetic endpoints into real entities when required by the relationship structure.

When `merge_entities=True`, the `EntityResolver` class (implemented in [`semantica/kg/entity_resolver.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/entity_resolver.py)) merges duplicate entities using configurable strategies: `fuzzy`, `exact`, or `ml-based`. Following deduplication, `GraphBuilder._remap_relationship_endpoints` (lines 66-86) rewrites relationship source and target IDs to match the canonical identifiers produced by the resolver.

## Conflict Detection and Graph Assembly

An optional `ConflictDetector` (referenced at lines 57-66 in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py) and implemented in [`semantica/conflicts/conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/conflict_detector.py)) scans entity lists for contradictory information and attempts automated resolution when enabled.

The core `GraphBuilder.build` method (lines 18-27 and 64-78) assembles the final graph dictionary containing:

- `entities`: The resolved, deduplicated node list
- `relationships`: The remapped edge list
- `metadata`: Statistics including counts, timestamps, and temporal feature flags

## Persistence and Temporal Extensions

When a `GraphStore` instance is supplied—such as `Neo4jGraphStore` from [`semantica/graph_store/neo4j_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/neo4j_store.py)—the framework automatically persists nodes and edges to external graph databases during the build process. Setting `enable_temporal=True` (lines 45-48) adds temporal validity to edges and enables version snapshots for time-aware queries.

```python
from semantica.graph_store.neo4j_store import Neo4jGraphStore
from semantica.kg import GraphBuilder

store = Neo4jGraphStore(uri="bolt://localhost:7687", auth=("neo4j", "password"))
builder = GraphBuilder(graph_store=store, enable_temporal=True)
builder.build("Your unstructured text here...")  # Auto-persists to Neo4j

```

## Summary

- The pipeline starts with unified ingestion via `ingest()` and specialized Ingestor classes defined in `semantica/ingest/`.
- `GraphBuilder._extract_from_text` utilizes **NERExtractor**, **RelationExtractor**, and **TripletExtractor** with local `ml`, `pattern`, or `llm` backends.
- `_process_item` normalizes all inputs into canonical entity and relationship dictionaries.
- **EntityResolver** deduplicates nodes using fuzzy, exact, or ML-based strategies when `merge_entities=True`.
- `_remap_relationship_endpoints` updates relationship references to canonical IDs after deduplication.
- Optional **ConflictDetector** resolves contradictory information before final assembly.
- `GraphBuilder.build()` produces a dictionary with `entities`, `relationships`, and `metadata`.
- Integration with **GraphStore** implementations enables persistence to databases like Neo4j, with optional temporal extensions.

## Frequently Asked Questions

### What types of unstructured data sources does Semantica support?

Semantica accepts files, web pages, feeds, streams, and databases through its unified ingestion layer. The `ingest` function auto-detects source types and delegates to specialized classes like `FileIngestor`, `WebIngestor`, and `PublicAPIIngestor` as documented in [`semantica/ingest/ingest_usage.md`](https://github.com/semantica-agi/semantica/blob/main/semantica/ingest/ingest_usage.md).

### Does Semantica require external APIs for knowledge graph construction?

No. By default, Semantica uses local extraction methods that require no external APIs. The **NERExtractor**, **RelationExtractor**, and **TripletExtractor** support configurable backends including `ml`, `pattern`, and `llm`, all of which can operate locally according to the implementation in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py).

### How does Semantica handle duplicate entities during graph construction?

When `merge_entities=True`, the framework invokes **EntityResolver** to merge duplicates using `fuzzy`, `exact`, or `ml-based` matching strategies. After resolution, `GraphBuilder._remap_relationship_endpoints` updates all relationship references to point to the canonical entity IDs, ensuring graph consistency.

### Can Semantica persist knowledge graphs to external databases?

Yes. By passing a **GraphStore** instance—such as `Neo4jGraphStore` from [`semantica/graph_store/neo4j_store.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/graph_store/neo4j_store.py)—to the `GraphBuilder` constructor, the framework automatically persists entities and relationships during the build process. The system also supports temporal extensions when `enable_temporal=True` is configured.