# How Semantica Detects and Handles Conflicts Before Merging Data

> Semantica detects and handles data conflicts before merging by grouping entities by URI and scanning for divergences. Resolve label conflicts automatically for cleaner data.

- Repository: [Semantica /semantica](https://github.com/semantica-agi/semantica)
- Tags: how-to-guide
- Published: 2026-09-12

---

**Semantica detects conflicts by grouping entities by canonical URI and scanning for divergent labels, properties, and structural relationships, then optionally resolves label conflicts automatically before committing data to the graph store.**

The semantica-agi/semantica repository implements a lightweight, stateless conflict detection pipeline that runs during the ingestion phase. According to the source code in [`semantica/conflicts/conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/conflict_detector.py), the system identifies three primary conflict types and surfaces them before the final merge operation, ensuring graph consistency and deterministic data integration.

## The Three-Stage Conflict Detection Pipeline

The `ConflictDetector` class orchestrates a three-stage algorithm that processes entities before they reach the `GraphStore`. This design keeps the core logic testable while allowing the `GraphBuilder` to inject resolution strategies.

### Stage 1: Entity Grouping by Canonical URI

All incoming `Entity` objects are first grouped by their `uri` attribute—the unique identifier for a logical concept. The helper function `_group_by_entity` in [`semantica/conflicts/conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/conflict_detector.py) aggregates records from different sources that reference the same entity, creating the comparison sets needed for conflict analysis.

Entities sharing the same URI but originating from different data sources are compared field-by-field to surface inconsistencies.

### Stage 2: Multi-Dimensional Conflict Identification

For each URI group, the `detect_conflicts` method examines three specific dimensions:

- **Label conflicts** – If grouped records contain different `label` values (e.g., "Person" vs. "Human"), the system creates a `Conflict` object with type `ConflictType.LABEL`.

- **Property conflicts** – The detector compares every attribute except `uri` and `label` across the group. When concrete values differ (missing values are ignored), it generates a `ConflictType.PROPERTY` instance containing the property name and divergent values.

- **Structural conflicts** – If an entity defines `required_relations` (specific relationship types that must exist) and none of the grouped records contain a required relation type, the system raises a `ConflictType.STRUCTURAL` conflict.

Each conflict is encapsulated in a `Conflict` dataclass containing the entity URI, human-readable details, and a structured payload for programmatic handling.

### Stage 3: Automated Resolution and Reporting

After detection, the `GraphBuilder` invokes `resolve_conflicts` from [`conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/conflict_detector.py). Currently, only **label conflicts** support automatic resolution through an injected `label_resolver` strategy. By default, this strategy selects the first label in the list, though callers can provide custom logic. Property and structural conflicts remain unresolved and are emitted as warnings for downstream tooling or manual inspection.

## Core Implementation in conflict_detector.py

The `ConflictDetector` class in [`semantica/conflicts/conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/conflict_detector.py) implements the detection logic as a stateless service:

```python
from semantica.conflicts.conflict_detector import ConflictDetector, ConflictType

detector = ConflictDetector(label_resolver=lambda labels: labels[0])
conflicts = detector.detect_conflicts(entity_list)

```

The `detect_conflicts` method returns a list of `Conflict` objects, each serializable via `to_dict()` for CLI JSON output. The `resolve_conflicts` method processes these objects and returns a summary dictionary indicating how many conflicts were automatically fixed.

## Integration with GraphBuilder in graph_builder.py

The `GraphBuilder` class in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py) orchestrates the conflict detection workflow during the build phase. When initialized with `resolve_conflicts=True` (the default), the builder instantiates a `ConflictDetector` and runs it against entity batches before normalization and storage:

```python
from semantica.kg.graph_builder import GraphBuilder
from semantica.graph.store import GraphStore

store = GraphStore()
builder = GraphBuilder(store, resolve_conflicts=True)
merged = builder.build(entity_iterable)

```

The `build` method performs three conceptual steps:
1. **Conflict detection** – Runs the detector against the entity batch
2. **Resolution** – Automatically resolves label conflicts and logs counts
3. **Write** – Normalizes and persists the entities via `store.write()`

## Working with the ConflictDetector API

You can invoke conflict detection directly for debugging or CLI usage without instantiating `GraphBuilder`:

```python
from semantica.graph.core import Entity
from semantica.conflicts.conflict_detector import detect_conflicts, ConflictDetector

# Create conflicting entities

e1 = Entity(uri="ex:Person/1", label="Person", properties={"age": 30})
e2 = Entity(uri="ex:Person/1", label="Individual", properties={"age": 31})

# Method 1: Direct API (returns Conflict objects)

detector = ConflictDetector()
conflicts = detector.detect_conflicts([e1, e2])
for c in conflicts:
    print(f"{c.conflict_type}: {c.details}")

# Method 2: CLI wrapper (returns dicts)

json_ready = detect_conflicts([e1, e2])

```

The `Conflict` payload for property conflicts includes the property name and sorted values, while label conflicts include the candidate label list—enabling downstream services to implement custom resolution logic without importing the full Python class hierarchy.

## Summary

- **Conflict detection occurs before merging** – The `GraphBuilder` runs `ConflictDetector` immediately after entity collection but before calling `store.write()`.
- **Three conflict types are surfaced** – Label, property, and structural conflicts are detected by comparing entities grouped via canonical URI.
- **Stateless design** – The `ConflictDetector` class requires no persistent state, making it suitable for CLI tools, unit tests, and pipeline integration.
- **Pluggable resolution** – Only label conflicts support automatic resolution via the `label_resolver` parameter; other conflict types require manual intervention.
- **Full provenance** – Each `Conflict` object carries structured payloads suitable for JSON serialization and downstream audit systems.

## Frequently Asked Questions

### What types of conflicts does Semantica detect?

Semantica detects three conflict types defined in `ConflictType`: **LABEL** (differing classifications), **PROPERTY** (inconsistent attribute values), and **STRUCTURAL** (missing required relationships). The detection logic resides in [`semantica/conflicts/conflict_detector.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/conflicts/conflict_detector.py) and examines all entities sharing the same URI.

### How does Semantica resolve label conflicts?

Label conflicts are resolved using an optional **label resolver strategy** passed to `ConflictDetector.__init__`. The default implementation selects the first label from the candidate list, but callers can inject custom logic. The `resolve_conflicts` method updates its resolution count based on successfully handled label conflicts.

### Can structural and property conflicts be auto-resolved?

Currently, the built-in resolver only handles **LABEL** conflicts. Property and structural conflicts are detected and logged but left untouched by `resolve_conflicts`. Downstream systems must implement domain-specific policies to resolve these conflict types manually or extend the `ConflictDetector` with additional resolution strategies.

### Where does conflict detection occur in the ingestion pipeline?

Conflict detection runs inside `GraphBuilder.build()` in [`semantica/kg/graph_builder.py`](https://github.com/semantica-agi/semantica/blob/main/semantica/kg/graph_builder.py), positioned between entity normalization and the final write operation to the `GraphStore`. This placement ensures that contradictory information never reaches the persistent storage layer when `resolve_conflicts=True`.