How to Choose Between Deidentification Methods (Mask, Replace, Hash) in OpenMED
Choose mask for deterministic placeholder redaction, replace for realistic synthetic data, and hash for stable opaque identifiers that enable cross-record linking without exposing PII.
OpenMED's deidentification pipeline, implemented in the maziyarpanahi/openmed repository, provides a unified API for redacting sensitive information using multiple strategies. The primary entry point openmed.core.pii.deidentify orchestrates language-aware PII detection and delegates to _redact_entity to apply your chosen transformation, allowing you to balance privacy compliance with downstream data utility.
How OpenMED Deidentification Works
The deidentification process follows a consistent pipeline regardless of which method you select. First, deidentify (defined in openmed/core/pii.py, lines 752–770) extracts PII entities using language-specific detection models. Then, the internal _redact_entity function (lines 558–610) applies a switch-statement to transform each detected entity according to the method parameter. Finally, the function aggregates results into a DeidentificationResult object that optionally preserves a mapping for re-identification.
All methods share the same detection backbone, ensuring that language-specific models (lang argument) and confidence thresholds are respected uniformly across your pipeline.
The Five Deidentification Strategies Explained
OpenMED supports five distinct redaction strategies, each optimized for specific privacy and utility requirements.
Mask (Placeholder Redaction)
The mask method replaces each detected entity with a generic placeholder indicating its type, such as [NAME], [EMAIL], or [PHONE]. According to the source code in openmed/core/pii.py (line 580), this returns f"[{entity.entity_type}]".
When to use: Select this method for audit logs, model training data where token positions must be preserved, or any scenario requiring deterministic, human-readable indicators of where PII was removed.
Remove (Complete Deletion)
The remove strategy deletes the entity entirely, replacing it with an empty string (line 584). This reduces text length and eliminates any trace of the original entity's position.
When to use: Choose this for storage-constrained pipelines or when the exact location of PII is irrelevant to downstream processing.
Replace (Synthetic Surrogates)
The replace method generates realistic-looking surrogate data using the faker library or a custom Anonymizer. As implemented in lines 888–896, this calls _generate_fake_pii or anonymizer.surrogate to create contextually appropriate replacements (e.g., "John Doe" becomes "Sarah Smith").
When to use: Use this for data-sharing scenarios where text must remain readable and natural, such as training NLP models. Set consistent=True with a fixed seed to ensure the same original value maps to the same surrogate across your entire dataset, enabling reproducible anonymization.
Hash (Deterministic Tokenization)
The hash method produces a short, deterministic hash like NAME_1a2b3c4d that uniquely identifies the original value without revealing it. The implementation (lines 998–1002) computes a SHA-256 hash of the entity text and truncates it to 8 hexadecimal characters.
When to use: Select this when you need to track entities across multiple documents or records (e.g., longitudinal patient studies) without exposing actual identifiers. The hash acts as a stable opaque key for analytics while maintaining cryptographic separation from the original PII.
Shift Dates (Temporal Obfuscation)
The shift_dates method shifts every detected date by a fixed offset, preserving temporal intervals between dates while anonymizing absolute values. Non-date entities fall back to masking. This delegates to _shift_date (lines 1004–1008).
When to use: Use this for HIPAA-compliant date obfuscation where relative timelines (e.g., "30 days later") must be preserved for clinical analysis. Set keep_year=False to fully randomize the year component when needed.
Decision Framework: When to Use Each Method
Your choice depends on whether you prioritize compliance guarantees, data utility, or entity linkability:
- Compliance-first scenarios: Use
maskorremovefor the safest defaults. These methods never emit synthetic data that could be mistaken for real identifiers, ensuring zero risk of PII leakage through surrogate generation. - Data utility requirements: Use
replacewhen downstream models benefit from realistic context. This method preserves text flow and readability while protecting privacy. Remember to useconsistent=Trueand a fixedseedwhen you need deterministic mapping across batches. - Cross-record analytics: Use
hashwhen you must join records on the same hidden identifier. This provides a stable token for entity resolution without revealing the original value, making it ideal for patient-level analytics across multiple encounters. - Temporal data preservation: Use
shift_datesfor clinical timelines or cohort studies where date intervals carry medical significance but absolute dates must be protected.
Implementation Examples
The following examples demonstrate how to invoke each method using the deidentify function. These snippets assume OpenMED is installed with the default English model available.
Basic Masking (Default)
from openmed.core.pii import deidentify
sample = "Patient John Doe (DOB: 01/15/1970) called from 555-1234"
masked = deidentify(sample, method="mask")
print(masked.deidentified_text)
# → Patient [NAME] (DOB: [DATE]/1970) called from [PHONE]
Realistic Replacement with Consistency
# Generate deterministic surrogates using a fixed seed
replaced = deidentify(
sample,
method="replace",
lang="en",
consistent=True,
seed=42
)
print(replaced.deidentified_text)
# → Patient Johnathan Smith (DOB: 02/14/2020) called from 555-9876
Batch Processing with Consistent Surrogates
texts = [
"Patient John Doe visited Dr. Alice Smith.",
"John Doe reported improvement."
]
# Consistent mapping across the batch
batch = [deidentify(t, method="replace", consistent=True, seed=99) for t in texts]
for r in batch:
print(r.deidentified_text)
# Ensures "John Doe" maps to the same surrogate in both texts
Linkable Hashing for Analytics
hashed = deidentify(sample, method="hash")
print(hashed.deidentified_text)
# → Patient NAME_5a1b2c3d (DOB: DATE_9f8e7d6c)/1970 called from PHONE_4d3c2b1a
# Extract hash values for database joins
hash_map = {e.entity_type: e.hash_value for e in hashed.pii_entities}
Date Shifting with Interval Preservation
shifted = deidentify(
"Patient was admitted on 01/15/2020 and discharged on 02/01/2020.",
method="shift_dates",
keep_year=False, # Fully randomize the year
date_shift_days=30 # Shift all dates by 30 days
)
print(shifted.deidentified_text)
# → Patient was admitted on 02/14/2021 and discharged on 03/02/2021.
# The 17-day interval between admission and discharge is preserved.
Summary
- Mask provides deterministic placeholders like
[NAME]when you need to preserve token positions and indicate where redaction occurred. - Replace generates realistic synthetic data using
faker, ideal for maintaining text utility in NLP training datasets; useconsistent=True+seedfor reproducible mappings. - Hash creates stable SHA-256-based tokens (truncated to 8 characters) that enable cross-document entity linking without exposing original values.
- Remove deletes entities entirely for maximum space efficiency when entity positions are irrelevant.
- Shift Dates preserves temporal intervals while obfuscating absolute dates, essential for HIPAA-compliant clinical timelines.
All methods are implemented in openmed/core/pii.py and validated in tests/unit/test_pii.py (see test_deidentify_mask_method, test_deidentify_replace_method, and test_deidentify_hash_method).
Frequently Asked Questions
Can I use hash for dates or only for text entities?
The hash method works only for the entity's original text, not for dates. According to the implementation in openmed/core/pii.py, date entities have specific handling requirements and should use the shift_dates method instead. If you need to obfuscate dates while maintaining the ability to link records temporally, use shift_dates with a consistent date_shift_days value across your dataset.
Is the replace method deterministic or random?
By default, replace uses the faker library to generate random surrogates. However, you can make it deterministic by setting consistent=True and providing a fixed seed parameter. When configured this way, the same input value will always produce the same surrogate output across multiple calls and documents, which is essential for reproducible data pipelines and consistent anonymization.
How secure is the hash method for protecting PII?
The hash method uses SHA-256 truncated to 8 hexadecimal characters, which provides a one-way transformation suitable for linking records but not for cryptographically secure storage of sensitive identifiers. While adequate for analytics and internal entity resolution, note that short hashes may be vulnerable to brute-force attacks if the original value space is small (e.g., short names). For cryptographically secure requirements, consider using mask or remove instead.
What happens if I select shift_dates but the text contains names and phone numbers?
When using method="shift_dates", the system applies date shifting only to detected date entities. All other PII types (names, phone numbers, emails, etc.) automatically fall back to the mask behavior. This hybrid approach ensures that dates are handled with temporal preservation while other sensitive data receives standard placeholder redaction, as implemented in the _redact_entity switch statement (lines 558–610).
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →