Semantica's Strategies for Conflict Detection and Resolution: A Technical Deep Dive
Semantica implements a comprehensive conflicts sub-package that provides seven distinct detection methods and seven resolution strategies to identify and reconcile inconsistencies across multi-source knowledge graphs.
The semantica-agi/semantica repository embeds a production-grade pipeline for conflict detection and resolution that operates on heterogeneous entity data. By combining value-based comparison algorithms with credibility-weighted resolution logic, Semantica enables both automated reconciliation and human-in-the-loop review workflows.
Conflict Detection Strategies
The ConflictDetector class in semantica/conflicts/conflict_detector.py exposes specialized methods for identifying specific categories of data inconsistencies. All detectors return standardized Conflict dataclass instances containing conflict ID, type, entity ID, property details, conflicting values, source provenance, confidence scores, and severity levels.
Value-Based and Property-Wide Detection
The detect_value_conflicts method groups entities by ID and flags conflicts when len(set(values)) > 1 for a given property. This core function identifies cases where the same entity ID reports differing values across sources (e.g., two databases listing different company names).
The detect_property_conflicts method serves as a thin wrapper around the value detector, allowing arbitrary property inspection. For comprehensive entity scanning, detect_entity_conflicts dynamically builds field lists (excluding meta-fields) and iterates over all properties using the value detector logic.
Relationship and Type Conflict Detection
Relationship inconsistencies are captured by detect_relationship_conflicts, which groups relationships by composite keys and checks each attribute for divergent values across sources. This detects conflicting relationship types, properties, or confidence scores.
Type conflicts are handled separately via detect_type_conflicts, which applies similar grouping logic but focuses exclusively on the type field. This catches cases where the same entity is classified as both "Person" and "Organization" across different knowledge sources.
Temporal and Logical Conflict Detection
The detect_temporal_conflicts method searches for predefined temporal property names, normalizes values (extracting years from date strings), and flags entities with differing normalized temporal data (e.g., founded year 1998 versus 2003).
Logical inconsistencies are identified by detect_logical_conflicts, which uses a hard-coded incompatibility map to detect prohibited type combinations (e.g., an entity simultaneously classified as "Person" and "Organization"). This method raises conflicts when incompatible type pairs appear within the same entity's sources.
Unified Detection Dispatcher
The detect_conflicts method acts as the central entry point, accepting a method argument ("value", "property", "type", "relationship", "temporal", "logical", or "all") and routing to the appropriate detector. All detection methods integrate with get_progress_tracker() for observable processing of large datasets.
from semantica.conflicts import ConflictDetector
detector = ConflictDetector()
all_conflicts = detector.detect_conflicts(graph, method="all")
print(f"Found {len(all_conflicts)} conflicts")
Conflict Resolution Strategies
The ConflictResolver class in semantica/conflicts/conflict_resolver.py implements the ResolutionStrategy enum, providing seven distinct resolution algorithms that can be applied globally or per-property.
Voting and Credibility-Weighted Resolution
Voting (VOTING) resolves conflicts by selecting the most frequent value among candidates, calculating confidence as the vote ratio. This simple majority approach works effectively when source reliability is uniform.
Credibility-Weighted resolution (CREDIBILITY_WEIGHTED) enhances the voting mechanism by weighting each candidate according to source credibility (retrieved from SourceTracker) and reported confidence scores. The implementation in _resolve_by_credibility selects the highest-weight value, prioritizing trustworthy sources over raw frequency.
Temporal and Confidence-Based Resolution
Most Recent (MOST_RECENT) resolution selects the value from the source bearing the newest timestamp in metadata["timestamp"]. This strategy prioritizes data freshness when temporal accuracy correlates with correctness.
Highest Confidence (HIGHEST_CONFIDENCE) resolution directly compares confidence scores reported by sources and selects the maximum value. This works effectively when sources provide calibrated confidence estimates.
First Seen (FIRST_SEEN) resolution maintains source ordering by selecting the earliest appearing value in the source list, preserving original data precedence.
Manual and Expert Review Workflows
Manual Review (MANUAL_REVIEW) flags conflicts for human inspection by returning a ResolutionResult with resolved=False and appropriate metadata markers. Similarly, Expert Review (EXPERT_REVIEW) signals that specialized domain knowledge is required, routing conflicts to subject matter experts rather than general reviewers.
Configurable Resolution Rules
The resolver supports fine-grained control through default_strategy configuration and per-entity/property rules managed via set_resolution_rule. When resolve_conflict or resolve_conflicts is called, the system normalizes strategy selection (falling back to defaults or property-specific rules), dispatches to the appropriate private method, and records ResolutionResult instances in an internal history log accessible via get_resolution_history().
from semantica.conflicts import ConflictResolver, ResolutionStrategy
resolver = ConflictResolver(default_strategy=ResolutionStrategy.VOTING)
resolution_results = resolver.resolve_conflicts(all_conflicts, strategy="voting")
# Configure custom rules per property
resolver.set_resolution_rule(
entity_id="company_123",
property_name="revenue",
strategy="credibility_weighted"
)
End-to-End Implementation Workflow
Implementing conflict detection and resolution in Semantica follows a four-stage pipeline:
-
Detect – Execute detection methods on knowledge graphs or entity collections using
ConflictDetector. -
Report – Generate structured analytics via
ConflictDetector.get_conflict_report(), which returns counts by type/severity and detailed conflict entries. -
Resolve – Process
Conflictobjects throughConflictResolver.resolve_conflicts(), optionally specifying strategies like"voting"or"most_recent". -
Audit – Inspect
ResolutionResultobjects and resolution history to verify changes and maintain provenance tracking.
# Generate human-readable conflict reports
report = detector.get_conflict_report()
print(report["total_conflicts"])
print(report["by_type"])
print(report["by_severity"])
Summary
-
Semantica provides seven detection methods (value, property, relationship, type, temporal, logical, and unified dispatch) through
semantica/conflicts/conflict_detector.py. -
The seven resolution strategies (voting, credibility-weighted, most recent, first seen, highest confidence, manual review, and expert review) are implemented in
semantica/conflicts/conflict_resolver.py. -
All detection methods return standardized
Conflictdataclass instances containing provenance, confidence, and severity metadata. -
Resolution supports both global defaults and granular per-entity/property rules via
set_resolution_rule(). -
Progress tracking integration ensures observable processing for large-scale knowledge graph reconciliation.
Frequently Asked Questions
How does Semantica detect conflicting values across different data sources?
Semantica uses the detect_value_conflicts method in semantica/conflicts/conflict_detector.py to group entities by ID and compare property values. When the set of values for a given property contains more than one distinct entry, the system generates a Conflict instance capturing the differing values, their respective sources, and confidence scores.
Can Semantica automatically resolve conflicts without human intervention?
Yes. The ConflictResolver class provides five automated strategies: Voting, Credibility-Weighted, Most Recent, First Seen, and Highest Confidence. These algorithms select winning values based on frequency, source trustworthiness, timestamps, or confidence scores. However, for critical conflicts, the system supports Manual Review and Expert Review strategies that flag items for human validation.
What is the difference between type conflicts and logical conflicts in Semantica?
Type conflicts (detected by detect_type_conflicts) occur when the same entity ID is assigned different type values across sources (e.g., "Person" versus "Organization"). Logical conflicts (detected by detect_logical_conflicts) use a hard-coded incompatibility map to identify semantically impossible type combinations, such as an entity simultaneously classified as both "Person" and "Organization" within the same merged dataset.
How can I configure different resolution strategies for specific entity properties?
Use the set_resolution_rule method on a ConflictResolver instance to assign strategy overrides for specific entity-property combinations. For example, you can apply credibility_weighted resolution to revenue fields while using most_recent for contact information. The resolver automatically applies these rules during resolve_conflicts() execution, falling back to the default_strategy for unspecified properties.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →