How Semantica Detects and Resolves Conflicts Before Deduplication

Semantica resolves contradictory information from multiple sources by running a ConflictDetector to identify value, type, temporal, and logical clashes, then applies a ConflictResolver with configurable strategies (consensus, weighted, or rule-based) to normalize entities before the deduplication engine merges duplicate records.

The semantica-agi/semantica repository implements a robust knowledge graph pipeline that prevents ambiguous data from polluting the final graph. Before merging duplicate entities, the system actively detects and reconciles conflicting information across data sources. This pre-deduplication workflow ensures that only consistent, authoritative representations survive the deduplication step.

The Pre-Deduplication Conflict Pipeline

Semantica’s architecture separates conflict management into distinct detection and resolution phases that execute prior to entity merging. This sequencing prevents contradictory data from being collapsed into a single node during deduplication.

Conflict Detection Engine

The ConflictDetector class in semantica/conflicts/conflict_detector.py serves as the primary engine for identifying inconsistencies. During graph construction, the GraphBuilder instantiates this detector (see line 159 in semantica/kg/graph_builder.py) to scan incoming records grouped by entity or relationship identifiers.

The detector runs a multi-method analysis to surface specific inconsistency types:

  • Value conflicts – The detect_value_conflicts method examines a specific property across sources and flags a conflict when the value set contains more than one distinct entry (lines 71‑78 in conflict_detector.py).
  • Type conflicts – The detect_type_conflicts method ensures the same entity does not appear with incompatible semantic types (lines 66‑84).
  • Temporal conflicts – The detect_temporal_conflicts method identifies divergent timestamps or founding years that cannot be reconciled chronologically (lines 40‑47).
  • Logical conflicts – The detect_logical_conflicts method applies domain-specific rules, such as validating that a Person cannot simultaneously be an Organization (lines 48‑55).

Each detection routine generates a Conflict dataclass instance that records the conflict ID, type, involved entity, conflicting values, provenance sources, confidence scores, severity levels, and recommended actions (lines 78‑92).

Resolution Strategies

After detection, the pipeline hands the list of Conflict objects to the ConflictResolver (instantiated via ConflictResolver(**kwargs) in semantica/conflicts/methods.py, lines 121‑124). The resolver in semantica/conflicts/conflict_resolver.py implements three distinct strategies:

  • Consensus – Selects the value that appears most frequently across contributing sources.
  • Weighted – Assigns credibility weights to each source and selects the highest-confidence value.
  • Rule-based – Applies domain-specific logic, such as preferring the most recent timestamp or selecting from an authoritative source hierarchy.

The resolver returns a normalized entity set where each conflict has been addressed according to the chosen strategy.

Integration with Graph Construction

The GraphBuilder orchestrates the entire workflow. It first invokes the ConflictDetector to surface issues, optionally routes conflicts through the ConflictResolver to produce a cleaned dataset, and finally executes deduplicate_entities to merge records sharing canonical identifiers. Because all contradictory information has been reconciled during the resolution phase, the deduplication step safely collapses duplicates without propagating ambiguous data.

Implementation Example

The following workflow demonstrates how to manually execute detection and resolution before passing data to the graph builder:

from semantica.conflicts import ConflictDetector, ConflictResolver

# 1️⃣ Detect conflicts in raw entity dictionaries

detector = ConflictDetector()
value_conflicts = detector.detect_value_conflicts(entities, property_name="name")
type_conflicts = detector.detect_type_conflicts(entities)

# 2️⃣ Resolve using default consensus strategy

resolver = ConflictResolver()
resolved_entities = resolver.resolve(
    entities, 
    conflicts=value_conflicts + type_conflicts
)

# 3️⃣ Build graph with pre-resolved conflicts

from semantica.kg.graph_builder import GraphBuilder
builder = GraphBuilder(resolve_conflicts=False)  # Skip redundant resolution

graph = builder.build(resolved_entities)         # Deduplication occurs here

Key Files and Architecture

Component File Path Primary Role
Conflict Detection Engine semantica/conflicts/conflict_detector.py Scans entities for value, type, temporal, and logical clashes; defines the Conflict dataclass.
Conflict Resolution Engine semantica/conflicts/conflict_resolver.py Implements consensus, weighted, and rule-based strategies to select single representations.
Resolution Facade semantica/conflicts/methods.py Provides the resolve_conflicts convenience function that wires detector to resolver (lines 121‑124).
Graph Construction semantica/kg/graph_builder.py Instantiates detection, runs resolution, and executes deduplication (line 159).
Provenance Tracking semantica/conflicts/source_tracker.py Records which source contributed each conflicting value for attribution during resolution.

Summary

  • Detection precedes merging – The ConflictDetector identifies value, type, temporal, and logical inconsistencies before any deduplication occurs.
  • Configurable resolution – The ConflictResolver supports consensus, weighted credibility, and rule-based strategies to normalize conflicting data.
  • Pipeline integration – GraphBuilder orchestrates the sequence: detect → resolve → deduplicate, ensuring clean data enters the knowledge graph.
  • Provenance awareness – Every conflict retains source attribution via source_tracker.py, enabling traceable resolution decisions.

Frequently Asked Questions

What types of conflicts does Semantica detect before deduplication?

Semantica detects four primary conflict categories: value conflicts (contradictory property values), type conflicts (incompatible semantic classifications), temporal conflicts (irreconcilable timestamps or dates), and logical conflicts (domain-rule violations such as a entity being both a Person and Organization). Each type is identified by dedicated methods in conflict_detector.py.

How does the ConflictResolver choose which value to keep?

The resolver’s choice depends on the active strategy. The consensus strategy selects the most frequently occurring value, the weighted strategy calculates a confidence score across sources and picks the highest, and the rule-based strategy applies custom logic such as preferring the most recent timestamp or a specific authoritative source.

Why does conflict resolution occur before deduplication rather than after?

Resolving conflicts before deduplication prevents contradictory information from being merged into a single canonical entity. If deduplication occurred first, conflicting values from different sources would be collapsed into one node, creating data integrity issues that would be harder to trace and fix retroactively.

How does Semantica track the source of conflicting information?

The system uses the SourceTracker class in semantica/conflicts/source_tracker.py to record provenance for every extracted fact. When the ConflictDetector creates a Conflict instance, it embeds these source references, allowing the ConflictResolver to apply source-specific weights and maintain audit trails for every resolution decision.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →