How Semantica Transforms Unstructured Data into a Knowledge Graph

Semantica converts unstructured data into structured knowledge graphs through a multi-stage pipeline that includes ingestion, entity extraction, normalization, deduplication, and graph assembly, implemented primarily in the GraphBuilder class.

The open-source Semantica framework (semantica-agi/semantica) provides a robust engine to transform unstructured data into a knowledge graph without requiring external APIs by default. By orchestrating specialized ingestors, local extractors, and resolution algorithms, Semantica converts raw text, files, and web sources into query-ready graph structures.

Unified Ingestion Across Heterogeneous Sources

The pipeline begins with the public ingest function, which auto-detects source types and routes data to the appropriate Ingestor class. According to semantica/ingest/ingest_usage.md (lines 30-45), the system supports FileIngestor, WebIngestor, PublicAPIIngestor, and other specialized handlers to process files, web pages, feeds, streams, and databases.

from semantica.ingest import ingest

# Auto-detect and ingest a PDF file

result = ingest("reports/annual_report.pdf", source_type="file")

Multi-Strategy Text Extraction

For textual sources, GraphBuilder._extract_from_text in semantica/kg/graph_builder.py (lines 48-66) orchestrates three local extractors:

  • NERExtractor identifies named entities
  • RelationExtractor extracts explicit relationships (optional)
  • TripletExtractor extracts subject-predicate-object triplets

All three support configurable backends—ml, pattern, or llm—and operate without external API dependencies by default.

from semantica.kg import GraphBuilder

builder = GraphBuilder(merge_entities=True, resolve_conflicts=True)
graph = builder.build(
    "The Acme Corp acquired Beta Ltd in 2023. Acme's CEO is Jane Doe.",
    extract=True,
    ner_method="ml",
    extract_relations=False
)

Normalization and Entity Resolution

The _process_item helper (lines 68-118 in semantica/kg/graph_builder.py) normalizes inputs—whether strings, Entity objects, Relation objects, or raw dictionaries—into canonical entity or relationship dictionaries. It also promotes synthetic endpoints into real entities when required by the relationship structure.

When merge_entities=True, the EntityResolver class (implemented in semantica/kg/entity_resolver.py) merges duplicate entities using configurable strategies: fuzzy, exact, or ml-based. Following deduplication, GraphBuilder._remap_relationship_endpoints (lines 66-86) rewrites relationship source and target IDs to match the canonical identifiers produced by the resolver.

Conflict Detection and Graph Assembly

An optional ConflictDetector (referenced at lines 57-66 in semantica/kg/graph_builder.py and implemented in semantica/conflicts/conflict_detector.py) scans entity lists for contradictory information and attempts automated resolution when enabled.

The core GraphBuilder.build method (lines 18-27 and 64-78) assembles the final graph dictionary containing:

  • entities: The resolved, deduplicated node list
  • relationships: The remapped edge list
  • metadata: Statistics including counts, timestamps, and temporal feature flags

Persistence and Temporal Extensions

When a GraphStore instance is supplied—such as Neo4jGraphStore from semantica/graph_store/neo4j_store.py—the framework automatically persists nodes and edges to external graph databases during the build process. Setting enable_temporal=True (lines 45-48) adds temporal validity to edges and enables version snapshots for time-aware queries.

from semantica.graph_store.neo4j_store import Neo4jGraphStore
from semantica.kg import GraphBuilder

store = Neo4jGraphStore(uri="bolt://localhost:7687", auth=("neo4j", "password"))
builder = GraphBuilder(graph_store=store, enable_temporal=True)
builder.build("Your unstructured text here...")  # Auto-persists to Neo4j

Summary

  • The pipeline starts with unified ingestion via ingest() and specialized Ingestor classes defined in semantica/ingest/.
  • GraphBuilder._extract_from_text utilizes NERExtractor, RelationExtractor, and TripletExtractor with local ml, pattern, or llm backends.
  • _process_item normalizes all inputs into canonical entity and relationship dictionaries.
  • EntityResolver deduplicates nodes using fuzzy, exact, or ML-based strategies when merge_entities=True.
  • _remap_relationship_endpoints updates relationship references to canonical IDs after deduplication.
  • Optional ConflictDetector resolves contradictory information before final assembly.
  • GraphBuilder.build() produces a dictionary with entities, relationships, and metadata.
  • Integration with GraphStore implementations enables persistence to databases like Neo4j, with optional temporal extensions.

Frequently Asked Questions

What types of unstructured data sources does Semantica support?

Semantica accepts files, web pages, feeds, streams, and databases through its unified ingestion layer. The ingest function auto-detects source types and delegates to specialized classes like FileIngestor, WebIngestor, and PublicAPIIngestor as documented in semantica/ingest/ingest_usage.md.

Does Semantica require external APIs for knowledge graph construction?

No. By default, Semantica uses local extraction methods that require no external APIs. The NERExtractor, RelationExtractor, and TripletExtractor support configurable backends including ml, pattern, and llm, all of which can operate locally according to the implementation in semantica/kg/graph_builder.py.

How does Semantica handle duplicate entities during graph construction?

When merge_entities=True, the framework invokes EntityResolver to merge duplicates using fuzzy, exact, or ml-based matching strategies. After resolution, GraphBuilder._remap_relationship_endpoints updates all relationship references to point to the canonical entity IDs, ensuring graph consistency.

Can Semantica persist knowledge graphs to external databases?

Yes. By passing a GraphStore instance—such as Neo4jGraphStore from semantica/graph_store/neo4j_store.py—to the GraphBuilder constructor, the framework automatically persists entities and relationships during the build process. The system also supports temporal extensions when enable_temporal=True is configured.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →