How Semantica Detects and Handles Conflicts Before Merging Data

Semantica detects conflicts by grouping entities by canonical URI and scanning for divergent labels, properties, and structural relationships, then optionally resolves label conflicts automatically before committing data to the graph store.

The semantica-agi/semantica repository implements a lightweight, stateless conflict detection pipeline that runs during the ingestion phase. According to the source code in semantica/conflicts/conflict_detector.py, the system identifies three primary conflict types and surfaces them before the final merge operation, ensuring graph consistency and deterministic data integration.

The Three-Stage Conflict Detection Pipeline

The ConflictDetector class orchestrates a three-stage algorithm that processes entities before they reach the GraphStore. This design keeps the core logic testable while allowing the GraphBuilder to inject resolution strategies.

Stage 1: Entity Grouping by Canonical URI

All incoming Entity objects are first grouped by their uri attribute—the unique identifier for a logical concept. The helper function _group_by_entity in semantica/conflicts/conflict_detector.py aggregates records from different sources that reference the same entity, creating the comparison sets needed for conflict analysis.

Entities sharing the same URI but originating from different data sources are compared field-by-field to surface inconsistencies.

Stage 2: Multi-Dimensional Conflict Identification

For each URI group, the detect_conflicts method examines three specific dimensions:

  • Label conflicts – If grouped records contain different label values (e.g., "Person" vs. "Human"), the system creates a Conflict object with type ConflictType.LABEL.

  • Property conflicts – The detector compares every attribute except uri and label across the group. When concrete values differ (missing values are ignored), it generates a ConflictType.PROPERTY instance containing the property name and divergent values.

  • Structural conflicts – If an entity defines required_relations (specific relationship types that must exist) and none of the grouped records contain a required relation type, the system raises a ConflictType.STRUCTURAL conflict.

Each conflict is encapsulated in a Conflict dataclass containing the entity URI, human-readable details, and a structured payload for programmatic handling.

Stage 3: Automated Resolution and Reporting

After detection, the GraphBuilder invokes resolve_conflicts from conflict_detector.py. Currently, only label conflicts support automatic resolution through an injected label_resolver strategy. By default, this strategy selects the first label in the list, though callers can provide custom logic. Property and structural conflicts remain unresolved and are emitted as warnings for downstream tooling or manual inspection.

Core Implementation in conflict_detector.py

The ConflictDetector class in semantica/conflicts/conflict_detector.py implements the detection logic as a stateless service:

from semantica.conflicts.conflict_detector import ConflictDetector, ConflictType

detector = ConflictDetector(label_resolver=lambda labels: labels[0])
conflicts = detector.detect_conflicts(entity_list)

The detect_conflicts method returns a list of Conflict objects, each serializable via to_dict() for CLI JSON output. The resolve_conflicts method processes these objects and returns a summary dictionary indicating how many conflicts were automatically fixed.

Integration with GraphBuilder in graph_builder.py

The GraphBuilder class in semantica/kg/graph_builder.py orchestrates the conflict detection workflow during the build phase. When initialized with resolve_conflicts=True (the default), the builder instantiates a ConflictDetector and runs it against entity batches before normalization and storage:

from semantica.kg.graph_builder import GraphBuilder
from semantica.graph.store import GraphStore

store = GraphStore()
builder = GraphBuilder(store, resolve_conflicts=True)
merged = builder.build(entity_iterable)

The build method performs three conceptual steps:

  1. Conflict detection – Runs the detector against the entity batch
  2. Resolution – Automatically resolves label conflicts and logs counts
  3. Write – Normalizes and persists the entities via store.write()

Working with the ConflictDetector API

You can invoke conflict detection directly for debugging or CLI usage without instantiating GraphBuilder:

from semantica.graph.core import Entity
from semantica.conflicts.conflict_detector import detect_conflicts, ConflictDetector

# Create conflicting entities

e1 = Entity(uri="ex:Person/1", label="Person", properties={"age": 30})
e2 = Entity(uri="ex:Person/1", label="Individual", properties={"age": 31})

# Method 1: Direct API (returns Conflict objects)

detector = ConflictDetector()
conflicts = detector.detect_conflicts([e1, e2])
for c in conflicts:
    print(f"{c.conflict_type}: {c.details}")

# Method 2: CLI wrapper (returns dicts)

json_ready = detect_conflicts([e1, e2])

The Conflict payload for property conflicts includes the property name and sorted values, while label conflicts include the candidate label list—enabling downstream services to implement custom resolution logic without importing the full Python class hierarchy.

Summary

  • Conflict detection occurs before merging – The GraphBuilder runs ConflictDetector immediately after entity collection but before calling store.write().
  • Three conflict types are surfaced – Label, property, and structural conflicts are detected by comparing entities grouped via canonical URI.
  • Stateless design – The ConflictDetector class requires no persistent state, making it suitable for CLI tools, unit tests, and pipeline integration.
  • Pluggable resolution – Only label conflicts support automatic resolution via the label_resolver parameter; other conflict types require manual intervention.
  • Full provenance – Each Conflict object carries structured payloads suitable for JSON serialization and downstream audit systems.

Frequently Asked Questions

What types of conflicts does Semantica detect?

Semantica detects three conflict types defined in ConflictType: LABEL (differing classifications), PROPERTY (inconsistent attribute values), and STRUCTURAL (missing required relationships). The detection logic resides in semantica/conflicts/conflict_detector.py and examines all entities sharing the same URI.

How does Semantica resolve label conflicts?

Label conflicts are resolved using an optional label resolver strategy passed to ConflictDetector.__init__. The default implementation selects the first label from the candidate list, but callers can inject custom logic. The resolve_conflicts method updates its resolution count based on successfully handled label conflicts.

Can structural and property conflicts be auto-resolved?

Currently, the built-in resolver only handles LABEL conflicts. Property and structural conflicts are detected and logged but left untouched by resolve_conflicts. Downstream systems must implement domain-specific policies to resolve these conflict types manually or extend the ConflictDetector with additional resolution strategies.

Where does conflict detection occur in the ingestion pipeline?

Conflict detection runs inside GraphBuilder.build() in semantica/kg/graph_builder.py, positioned between entity normalization and the final write operation to the GraphStore. This placement ensures that contradictory information never reaches the persistent storage layer when resolve_conflicts=True.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →