# How Smart Entity Merging Prevents PII Fragmentation in OpenMed

> Discover how smart entity merging in OpenMed stops PII fragmentation by reconstructing complete entities from fragmented data for better privacy protection.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**Smart entity merging prevents PII fragmentation by detecting complete semantic units via regex patterns, resolving overlapping token predictions, and blending pattern and model confidences to reconstruct single coherent entities from fragmented sub-tokens.**

Token-based NER models often split sensitive values like dates and SSNs into sub-tokens, causing fragmented predictions that break compliance workflows. The OpenMed library solves this through **smart entity merging**, a three-stage pipeline implemented in [`openmed/core/pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_entity_merger.py) that reconstructs complete PII units from partial model outputs. This approach ensures downstream redaction tools receive accurate, non-fragmented entities essential for privacy-preserving workflows.

## The PII Fragmentation Problem

When tokenizers process text like `"01/15/1970"`, they frequently split the date into separate tokens (e.g., `"01"` and `"/15/1970"`), causing the model to predict multiple partial entities rather than one complete **date-of-birth**. These fragments fail automated redaction systems and risk leaking protected information. Smart entity merging eliminates this issue by treating the full semantic unit as the atomic detection target.

## Three-Stage Smart Entity Merging

The `merge_entities_with_semantic_units` function in [`openmed/core/pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_entity_merger.py) implements a tightly coupled pipeline that eliminates fragmentation through three distinct stages.

### Stage 1: Semantic-Unit Detection with PIIPattern

The `PIIPattern` class scans raw text using regular expressions to identify whole PII units including dates, SSNs, and phone numbers. Each match receives a **context-aware confidence score** that is boosted when surrounding cue words are present and penalized if a custom validator fails (lines 468-511). This ensures that pattern-based detection accounts for textual context rather than relying solely on regex matches.

### Stage 2: Overlap Resolution and Dominant Label Selection

For every detected semantic unit, the merger gathers all model-predicted entities that spatially overlap it. The system selects the **dominant label** by averaging the model confidences or applying a tie-breaker when scores are equal (lines 600-613). This step aggregates fragmented sub-token predictions into a unified set of candidates for the full semantic span.

### Stage 3: Confidence Blending and Final Entity Creation

The final confidence score blends pattern and model contributions based on validation status:

- **Validated patterns**: 60% model confidence + 40% pattern confidence
- **Unvalidated patterns**: 90% model confidence + 10% pattern confidence (heavily discounted)

The final label selection depends on configuration flags `prefer_model_labels`, `allow_label_expansion`, and a specificity check via `is_more_specific` (lines 680-765). This logic determines whether to retain the model's label, the pattern's label, or the more specific of the two.

## Implementation: Reconstructing Fragmented Entities

The `merge_entities_with_semantic_units` function (lines 615-647) reconstructs complete entities from fragmented predictions. Here is how it processes a split date prediction:

```python
from openmed.core.pii_entity_merger import merge_entities_with_semantic_units

# Model predictions (often fragmented by tokenizer)

entities = [
    {"entity_type": "date", "score": 0.71, "start": 5, "end": 7, "word": "01"},
    {"entity_type": "date_of_birth", "score": 0.75, "start": 7, "end": 15, "word": "/15/1970"},
]

text = "DOB: 01/15/1970"

# Smart merging reconstructs the full entity

merged = merge_entities_with_semantic_units(entities, text)

print(merged[0])

# {'entity_type': 'date_of_birth',

#  'score': 0.80,

#  'start': 5,

#  'end': 15,

#  'word': '01/15/1970',

#  'merged_from': 2}

```

The function is called by [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py) (lines 386-407) during post-processing of model outputs, ensuring all PII entities are merged before downstream processing.

## Configuration Flags and Label Logic

The merger respects two key configuration flags that control label selection:

- **`prefer_model_labels`**: When enabled, the system prioritizes the model's predicted label over the regex pattern's label
- **`allow_label_expansion`**: Permits the merger to select a more specific label when `is_more_specific` determines that one candidate is a subclass of another (e.g., `date_of_birth` vs. `date`)

These flags work in conjunction with the confidence blending weights to determine whether the final entity carries the pattern's label, the model's label, or the most specific available label.

## Summary

- **Smart entity merging** reconstructs complete PII units from tokenizer fragments by detecting semantic units, resolving overlaps, and blending confidences
- The **`PIIPattern`** class provides regex-based detection with context-aware scoring in [`openmed/core/pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_entity_merger.py)
- **Confidence blending** uses 60/40 weights for validated patterns and 90/10 for unvalidated patterns
- Configuration flags **`prefer_model_labels`** and **`allow_label_expansion`** control final label selection
- Unit tests in [`tests/unit/test_pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_pii_entity_merger.py) verify that fragmentation is resolved for dates and other PII types

## Frequently Asked Questions

### What causes PII fragmentation in NER models?

PII fragmentation occurs when a tokenizer splits a single sensitive value (like a date or SSN) into multiple sub-tokens, causing the model to predict separate partial entities rather than one complete unit. This breaks entity boundaries and complicates redaction, as the system sees `"01"` and `"/15/1970"` instead of the full `"01/15/1970"`.

### How does OpenMed validate semantic units?

OpenMed validates semantic units through custom validators within the `PIIPattern` class that check the matched text against domain-specific rules. Validated patterns contribute 40% to the final confidence score and maintain full label authority, while unvalidated patterns are heavily discounted to only 10% contribution, preventing false positives from weak regex matches.

### What is the confidence blending formula?

The confidence blending formula depends on validation status. For validated patterns, the final score is **0.6 × model_confidence + 0.4 × pattern_confidence**. For unvalidated patterns, the pattern contribution is reduced to **0.9 × model_confidence + 0.1 × pattern_confidence**, ensuring that unreliable pattern matches do not override strong model predictions.

### Can I customize label preferences in the merger?

Yes, the merger supports two primary configuration flags: **`prefer_model_labels`**, which forces the system to retain the model's prediction even when a pattern suggests a different label, and **`allow_label_expansion`**, which permits the merger to upgrade to a more specific label when `is_more_specific` determines semantic subclass relationships between candidates.