How Semantica Handles Conflict Detection and Deduplication in Data Ingestion
Semantica uses a ConflictDetector class paired with the GraphBuilder to automatically identify contradictory statements and merge duplicate entities during ingestion, controlled via the resolve_conflicts and merge_entities boolean flags.
Semantica is an open-source knowledge graph construction framework that incrementally builds semantic graphs from streaming data sources. When ingesting new triples or entities, the system must reconcile contradictory values and collapse duplicate nodes to maintain graph integrity. This article examines the specific mechanisms for conflict detection and deduplication in data ingestion implemented in the semantica-agi/semantica repository.
The Core Components: ConflictDetector and GraphBuilder
The ingestion pipeline centers on two tightly coupled classes that handle data integrity.
ConflictDetector
In semantica/conflicts/conflict_detector.py, the ConflictDetector class scans incoming payloads for contradictory statements. When an entity already exists in the graph, the detector compares incoming predicate-value pairs against stored values. It exposes two primary methods: detect_conflicts(payload) which returns a list of conflict descriptors, and resolve_conflicts(conflicts, payload) which applies resolution strategies such as keep-first, keep-last, or custom merge logic.
GraphBuilder Orchestration
The GraphBuilder class in semantica/kg/graph_builder.py orchestrates the construction process. It accepts two critical boolean parameters: resolve_conflicts and merge_entities. When resolve_conflicts=True, the builder automatically instantiates and wires a ConflictDetector into the pipeline. Setting merge_entities=True enables deduplication logic that collapses nodes referring to the same real-world object before persistence.
The Conflict Detection Workflow
When resolve_conflicts is enabled, the ingestion process follows a four-stage pipeline:
- Normalization — Incoming records are converted into canonical triple format.
- Entity Lookup — The builder checks if the subject entity already exists in the graph.
- Conflict Detection — If the entity exists,
ConflictDetector.detect_conflicts()compares each incoming attribute against stored values. - Resolution — Depending on configuration, the system either rejects the contradictory statement or merges values using deterministic strategies like consensus voting.
Deduplication and Entity Merging
Deduplication occurs when GraphBuilder is initialized with merge_entities=True. The system treats records sharing the same identifier or fingerprint as a single node. During this process, duplicate nodes are collapsed and their outgoing edges are unified.
Crucially, the merging logic respects active conflict resolution settings. When both flags are enabled, contradictory attribute values are reconciled via the ConflictDetector before the final unified node is persisted to the graph.
Configuration Modes for Different Workloads
Semantica provides flexible ingestion modes through boolean flag combinations:
resolve_conflicts=True, merge_entities=True— Full integrity mode with conflict resolution and deduplication.resolve_conflicts=False, merge_entities=False— Append-only raw mode for maximum ingestion speed or debugging.resolve_conflicts=True, merge_entities=False— Conflict checking without deduplication, preserving duplicate entities but reconciling attribute contradictions.
Code Implementation Examples
The following examples demonstrate practical usage patterns for the conflict detection and deduplication API.
Standard ingestion with full conflict resolution and deduplication:
from semantica.kg.graph_builder import GraphBuilder
builder = GraphBuilder(resolve_conflicts=True, merge_entities=True)
graph = builder.build(triples_source) # Accepts lists, generators, or file paths
print(f"Graph contains {len(graph.nodes)} nodes")
High-speed raw ingestion without integrity checks:
builder = GraphBuilder(resolve_conflicts=False, merge_entities=False)
graph = builder.build(triples_source)
# Duplicate entities appear as separate nodes; no conflicts are detected
Manual conflict detection for custom workflows:
from semantica.conflicts.conflict_detector import ConflictDetector
detector = ConflictDetector()
conflicts = detector.detect_conflicts(incoming_payload)
resolved_payload = detector.resolve_conflicts(conflicts, incoming_payload)
# Resolved payload is ready for graph insertion
Key Source Files and References
The implementation spans several modules in the semantica-agi/semantica repository:
semantica/conflicts/conflict_detector.py— Core conflict detection and resolution logic.semantica/conflicts/conflicts_provenance.py— Provenance tracking for conflict resolution events.semantica/kg/graph_builder.py— Orchestrates ingestion with conflict and deduplication flags.cookbook/introduction/17_Conflict_Detection_and_Resolution.ipynb— Interactive tutorial for conflict handling.cookbook/introduction/18_Deduplication.ipynb— Practical guide to entity merging strategies.
Summary
- ConflictDetector provides the
detect_conflicts()andresolve_conflicts()methods for identifying and reconciling contradictory attribute values. - GraphBuilder controls ingestion behavior through
resolve_conflictsandmerge_entitiesboolean flags. - When both flags are enabled, the system normalizes data, checks for existing entities, detects conflicts, and merges duplicates while reconciling contradictions.
- Developers can disable these features for high-speed append-only ingestion or enable them for high-integrity knowledge graph construction.
- Source code resides in
semantica/conflicts/andsemantica/kg/directories, with accompanying notebooks in thecookbook/directory.
Frequently Asked Questions
What is the difference between conflict detection and deduplication in Semantica?
Conflict detection identifies when incoming data contradicts existing graph attributes, such as different values for the same property, while deduplication (merge_entities) collapses multiple records referring to the same real-world entity into a single node. These processes work together, with conflict resolution occurring before final node merging.
How do I disable conflict checking for faster data ingestion?
Initialize GraphBuilder with resolve_conflicts=False and merge_entities=False. This append-only mode skips all integrity checks, allowing duplicate entities and contradictory values to coexist in the graph for maximum ingestion throughput.
Can I use the ConflictDetector independently of GraphBuilder?
Yes. You can instantiate ConflictDetector directly from semantica.conflicts.conflict_detector, call detect_conflicts(payload) to identify issues, and use resolve_conflicts(conflicts, payload) to clean the data before manual graph insertion or external processing.
Where does Semantica track the history of resolved conflicts?
The semantica/conflicts/conflicts_provenance.py module handles provenance tracking, recording metadata about when conflicts were detected and which resolution strategy was applied, enabling audit trails for data lineage purposes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →