# How to Choose Between Deidentification Methods (Mask, Replace, Hash) in OpenMED

> Learn how to choose deidentification methods mask replace or hash in OpenMED. Select mask for placeholder redaction, replace for synthetic data, and hash for stable identifiers.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**Choose `mask` for deterministic placeholder redaction, `replace` for realistic synthetic data, and `hash` for stable opaque identifiers that enable cross-record linking without exposing PII.**

OpenMED's deidentification pipeline, implemented in the `maziyarpanahi/openmed` repository, provides a unified API for redacting sensitive information using multiple strategies. The primary entry point `openmed.core.pii.deidentify` orchestrates language-aware PII detection and delegates to `_redact_entity` to apply your chosen transformation, allowing you to balance privacy compliance with downstream data utility.

## How OpenMED Deidentification Works

The deidentification process follows a consistent pipeline regardless of which method you select. First, `deidentify` (defined in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py), lines 752–770) extracts PII entities using language-specific detection models. Then, the internal `_redact_entity` function (lines 558–610) applies a switch-statement to transform each detected entity according to the `method` parameter. Finally, the function aggregates results into a `DeidentificationResult` object that optionally preserves a mapping for re-identification.

All methods share the same detection backbone, ensuring that language-specific models (`lang` argument) and confidence thresholds are respected uniformly across your pipeline.

## The Five Deidentification Strategies Explained

OpenMED supports five distinct redaction strategies, each optimized for specific privacy and utility requirements.

### Mask (Placeholder Redaction)

The `mask` method replaces each detected entity with a generic placeholder indicating its type, such as `[NAME]`, `[EMAIL]`, or `[PHONE]`. According to the source code in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py) (line 580), this returns `f"[{entity.entity_type}]"`.

**When to use:** Select this method for audit logs, model training data where token positions must be preserved, or any scenario requiring deterministic, human-readable indicators of where PII was removed.

### Remove (Complete Deletion)

The `remove` strategy deletes the entity entirely, replacing it with an empty string (line 584). This reduces text length and eliminates any trace of the original entity's position.

**When to use:** Choose this for storage-constrained pipelines or when the exact location of PII is irrelevant to downstream processing.

### Replace (Synthetic Surrogates)

The `replace` method generates realistic-looking surrogate data using the `faker` library or a custom `Anonymizer`. As implemented in lines 888–896, this calls `_generate_fake_pii` or `anonymizer.surrogate` to create contextually appropriate replacements (e.g., "John Doe" becomes "Sarah Smith").

**When to use:** Use this for data-sharing scenarios where text must remain readable and natural, such as training NLP models. Set `consistent=True` with a fixed `seed` to ensure the same original value maps to the same surrogate across your entire dataset, enabling reproducible anonymization.

### Hash (Deterministic Tokenization)

The `hash` method produces a short, deterministic hash like `NAME_1a2b3c4d` that uniquely identifies the original value without revealing it. The implementation (lines 998–1002) computes a SHA-256 hash of the entity text and truncates it to 8 hexadecimal characters.

**When to use:** Select this when you need to **track** entities across multiple documents or records (e.g., longitudinal patient studies) without exposing actual identifiers. The hash acts as a stable opaque key for analytics while maintaining cryptographic separation from the original PII.

### Shift Dates (Temporal Obfuscation)

The `shift_dates` method shifts every detected date by a fixed offset, preserving temporal intervals between dates while anonymizing absolute values. Non-date entities fall back to masking. This delegates to `_shift_date` (lines 1004–1008).

**When to use:** Use this for HIPAA-compliant date obfuscation where relative timelines (e.g., "30 days later") must be preserved for clinical analysis. Set `keep_year=False` to fully randomize the year component when needed.

## Decision Framework: When to Use Each Method

Your choice depends on whether you prioritize compliance guarantees, data utility, or entity linkability:

- **Compliance-first scenarios:** Use `mask` or `remove` for the safest defaults. These methods never emit synthetic data that could be mistaken for real identifiers, ensuring zero risk of PII leakage through surrogate generation.
- **Data utility requirements:** Use `replace` when downstream models benefit from realistic context. This method preserves text flow and readability while protecting privacy. Remember to use `consistent=True` and a fixed `seed` when you need deterministic mapping across batches.
- **Cross-record analytics:** Use `hash` when you must join records on the same hidden identifier. This provides a stable token for entity resolution without revealing the original value, making it ideal for patient-level analytics across multiple encounters.
- **Temporal data preservation:** Use `shift_dates` for clinical timelines or cohort studies where date intervals carry medical significance but absolute dates must be protected.

## Implementation Examples

The following examples demonstrate how to invoke each method using the `deidentify` function. These snippets assume OpenMED is installed with the default English model available.

### Basic Masking (Default)

```python
from openmed.core.pii import deidentify

sample = "Patient John Doe (DOB: 01/15/1970) called from 555-1234"
masked = deidentify(sample, method="mask")
print(masked.deidentified_text)

# → Patient [NAME] (DOB: [DATE]/1970) called from [PHONE]

```

### Realistic Replacement with Consistency

```python

# Generate deterministic surrogates using a fixed seed

replaced = deidentify(
    sample, 
    method="replace", 
    lang="en", 
    consistent=True, 
    seed=42
)
print(replaced.deidentified_text)

# → Patient Johnathan Smith (DOB: 02/14/2020) called from 555-9876

```

### Batch Processing with Consistent Surrogates

```python
texts = [
    "Patient John Doe visited Dr. Alice Smith.",
    "John Doe reported improvement."
]

# Consistent mapping across the batch

batch = [deidentify(t, method="replace", consistent=True, seed=99) for t in texts]
for r in batch:
    print(r.deidentified_text)

# Ensures "John Doe" maps to the same surrogate in both texts

```

### Linkable Hashing for Analytics

```python
hashed = deidentify(sample, method="hash")
print(hashed.deidentified_text)

# → Patient NAME_5a1b2c3d (DOB: DATE_9f8e7d6c)/1970 called from PHONE_4d3c2b1a

# Extract hash values for database joins

hash_map = {e.entity_type: e.hash_value for e in hashed.pii_entities}

```

### Date Shifting with Interval Preservation

```python
shifted = deidentify(
    "Patient was admitted on 01/15/2020 and discharged on 02/01/2020.",
    method="shift_dates",
    keep_year=False,          # Fully randomize the year

    date_shift_days=30        # Shift all dates by 30 days

)
print(shifted.deidentified_text)

# → Patient was admitted on 02/14/2021 and discharged on 03/02/2021.

# The 17-day interval between admission and discharge is preserved.

```

## Summary

- **Mask** provides deterministic placeholders like `[NAME]` when you need to preserve token positions and indicate where redaction occurred.
- **Replace** generates realistic synthetic data using `faker`, ideal for maintaining text utility in NLP training datasets; use `consistent=True` + `seed` for reproducible mappings.
- **Hash** creates stable SHA-256-based tokens (truncated to 8 characters) that enable cross-document entity linking without exposing original values.
- **Remove** deletes entities entirely for maximum space efficiency when entity positions are irrelevant.
- **Shift Dates** preserves temporal intervals while obfuscating absolute dates, essential for HIPAA-compliant clinical timelines.

All methods are implemented in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py) and validated in [`tests/unit/test_pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_pii.py) (see `test_deidentify_mask_method`, `test_deidentify_replace_method`, and `test_deidentify_hash_method`).

## Frequently Asked Questions

### Can I use hash for dates or only for text entities?

The `hash` method works only for the entity's original text, not for dates. According to the implementation in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py), date entities have specific handling requirements and should use the `shift_dates` method instead. If you need to obfuscate dates while maintaining the ability to link records temporally, use `shift_dates` with a consistent `date_shift_days` value across your dataset.

### Is the replace method deterministic or random?

By default, `replace` uses the `faker` library to generate random surrogates. However, you can make it deterministic by setting `consistent=True` and providing a fixed `seed` parameter. When configured this way, the same input value will always produce the same surrogate output across multiple calls and documents, which is essential for reproducible data pipelines and consistent anonymization.

### How secure is the hash method for protecting PII?

The `hash` method uses SHA-256 truncated to 8 hexadecimal characters, which provides a one-way transformation suitable for linking records but not for cryptographically secure storage of sensitive identifiers. While adequate for analytics and internal entity resolution, note that short hashes may be vulnerable to brute-force attacks if the original value space is small (e.g., short names). For cryptographically secure requirements, consider using `mask` or `remove` instead.

### What happens if I select shift_dates but the text contains names and phone numbers?

When using `method="shift_dates"`, the system applies date shifting only to detected date entities. All other PII types (names, phone numbers, emails, etc.) automatically fall back to the `mask` behavior. This hybrid approach ensures that dates are handled with temporal preservation while other sensitive data receives standard placeholder redaction, as implemented in the `_redact_entity` switch statement (lines 558–610).