How Smart Entity Merging Prevents PII Fragmentation in OpenMed
Smart entity merging prevents PII fragmentation by detecting complete semantic units via regex patterns, resolving overlapping token predictions, and blending pattern and model confidences to reconstruct single coherent entities from fragmented sub-tokens.
Token-based NER models often split sensitive values like dates and SSNs into sub-tokens, causing fragmented predictions that break compliance workflows. The OpenMed library solves this through smart entity merging, a three-stage pipeline implemented in openmed/core/pii_entity_merger.py that reconstructs complete PII units from partial model outputs. This approach ensures downstream redaction tools receive accurate, non-fragmented entities essential for privacy-preserving workflows.
The PII Fragmentation Problem
When tokenizers process text like "01/15/1970", they frequently split the date into separate tokens (e.g., "01" and "/15/1970"), causing the model to predict multiple partial entities rather than one complete date-of-birth. These fragments fail automated redaction systems and risk leaking protected information. Smart entity merging eliminates this issue by treating the full semantic unit as the atomic detection target.
Three-Stage Smart Entity Merging
The merge_entities_with_semantic_units function in openmed/core/pii_entity_merger.py implements a tightly coupled pipeline that eliminates fragmentation through three distinct stages.
Stage 1: Semantic-Unit Detection with PIIPattern
The PIIPattern class scans raw text using regular expressions to identify whole PII units including dates, SSNs, and phone numbers. Each match receives a context-aware confidence score that is boosted when surrounding cue words are present and penalized if a custom validator fails (lines 468-511). This ensures that pattern-based detection accounts for textual context rather than relying solely on regex matches.
Stage 2: Overlap Resolution and Dominant Label Selection
For every detected semantic unit, the merger gathers all model-predicted entities that spatially overlap it. The system selects the dominant label by averaging the model confidences or applying a tie-breaker when scores are equal (lines 600-613). This step aggregates fragmented sub-token predictions into a unified set of candidates for the full semantic span.
Stage 3: Confidence Blending and Final Entity Creation
The final confidence score blends pattern and model contributions based on validation status:
- Validated patterns: 60% model confidence + 40% pattern confidence
- Unvalidated patterns: 90% model confidence + 10% pattern confidence (heavily discounted)
The final label selection depends on configuration flags prefer_model_labels, allow_label_expansion, and a specificity check via is_more_specific (lines 680-765). This logic determines whether to retain the model's label, the pattern's label, or the more specific of the two.
Implementation: Reconstructing Fragmented Entities
The merge_entities_with_semantic_units function (lines 615-647) reconstructs complete entities from fragmented predictions. Here is how it processes a split date prediction:
from openmed.core.pii_entity_merger import merge_entities_with_semantic_units
# Model predictions (often fragmented by tokenizer)
entities = [
{"entity_type": "date", "score": 0.71, "start": 5, "end": 7, "word": "01"},
{"entity_type": "date_of_birth", "score": 0.75, "start": 7, "end": 15, "word": "/15/1970"},
]
text = "DOB: 01/15/1970"
# Smart merging reconstructs the full entity
merged = merge_entities_with_semantic_units(entities, text)
print(merged[0])
# {'entity_type': 'date_of_birth',
# 'score': 0.80,
# 'start': 5,
# 'end': 15,
# 'word': '01/15/1970',
# 'merged_from': 2}
The function is called by openmed/core/pii.py (lines 386-407) during post-processing of model outputs, ensuring all PII entities are merged before downstream processing.
Configuration Flags and Label Logic
The merger respects two key configuration flags that control label selection:
prefer_model_labels: When enabled, the system prioritizes the model's predicted label over the regex pattern's labelallow_label_expansion: Permits the merger to select a more specific label whenis_more_specificdetermines that one candidate is a subclass of another (e.g.,date_of_birthvs.date)
These flags work in conjunction with the confidence blending weights to determine whether the final entity carries the pattern's label, the model's label, or the most specific available label.
Summary
- Smart entity merging reconstructs complete PII units from tokenizer fragments by detecting semantic units, resolving overlaps, and blending confidences
- The
PIIPatternclass provides regex-based detection with context-aware scoring inopenmed/core/pii_entity_merger.py - Confidence blending uses 60/40 weights for validated patterns and 90/10 for unvalidated patterns
- Configuration flags
prefer_model_labelsandallow_label_expansioncontrol final label selection - Unit tests in
tests/unit/test_pii_entity_merger.pyverify that fragmentation is resolved for dates and other PII types
Frequently Asked Questions
What causes PII fragmentation in NER models?
PII fragmentation occurs when a tokenizer splits a single sensitive value (like a date or SSN) into multiple sub-tokens, causing the model to predict separate partial entities rather than one complete unit. This breaks entity boundaries and complicates redaction, as the system sees "01" and "/15/1970" instead of the full "01/15/1970".
How does OpenMed validate semantic units?
OpenMed validates semantic units through custom validators within the PIIPattern class that check the matched text against domain-specific rules. Validated patterns contribute 40% to the final confidence score and maintain full label authority, while unvalidated patterns are heavily discounted to only 10% contribution, preventing false positives from weak regex matches.
What is the confidence blending formula?
The confidence blending formula depends on validation status. For validated patterns, the final score is 0.6 × model_confidence + 0.4 × pattern_confidence. For unvalidated patterns, the pattern contribution is reduced to 0.9 × model_confidence + 0.1 × pattern_confidence, ensuring that unreliable pattern matches do not override strong model predictions.
Can I customize label preferences in the merger?
Yes, the merger supports two primary configuration flags: prefer_model_labels, which forces the system to retain the model's prediction even when a pattern suggests a different label, and allow_label_expansion, which permits the merger to upgrade to a more specific label when is_more_specific determines semantic subclass relationships between candidates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →