Implementing PII De-identification with Replace Redaction in OpenMed

OpenMed's replace redaction strategy substitutes detected PII entities with realistic, locale-aware surrogates generated by the Faker library, supporting deterministic output through seeded hashing for reproducible de-identification across clinical datasets.

OpenMed provides a flexible pipeline for PII de-identification in clinical text, supporting multiple redaction strategies including mask, remove, replace, hash, and shift_dates. The replace method offers the most sophisticated approach by generating fake but realistic substitutes for sensitive entities like names, emails, and phone numbers. This article examines how to implement PII de-identification with replace redaction using the OpenMed source code, covering deterministic surrogate generation and multilingual support.

Understanding the Replace Redaction Strategy

The replace method stands among five available strategies in OpenMed. While mask obscures text with asterisks and remove deletes entities entirely, replace maintains text realism by injecting contextually appropriate fake data. This approach preserves the grammatical structure and semantic patterns of clinical narratives, making the de-identified text suitable for downstream NLP tasks and research datasets.

Locale-Aware Surrogate Generation

OpenMed leverages the Faker library to generate surrogates that match the linguistic context of the source document. When processing Portuguese clinical notes with lang="pt", the system produces Portuguese names and addresses rather than English equivalents. This locale awareness ensures that de-identified datasets maintain the cultural and linguistic characteristics necessary for training language-specific models.

Implementation Details in OpenMed

Entry Point and Core Logic

The de-identification pipeline centers on the deidentify function in openmed/core/pii.py. This function orchestrates NER model inference—using either Torch or MLX implementations—and delegates to specific redaction branches based on the method parameter. The function returns a DeidentificationResult object containing the processed text, surrogate mappings, and entity metadata.

The Replace Branch Implementation

Around line 930 in openmed/core/pii.py, the replace logic resides within an elif method == "replace": branch. For each detected PIIEntity, the code determines the appropriate Faker locale from the lang argument and generates type-specific surrogates by calling methods like fake.name(), fake.email(), or fake.phone_number(). Because surrogate lengths may differ from original entities, the algorithm updates the offset of subsequent entities to maintain correct character positions throughout the document.

Deterministic Replacement with Consistent Seeding

For research reproducibility, OpenMed supports deterministic surrogate generation through the consistent parameter. When consistent=True is passed alongside a seed value, the code hashes the original entity text and merges it with the user-provided seed. This hash initializes the Faker generator, ensuring that identical entities always map to identical surrogates across multiple runs. This implementation is crucial for longitudinal studies where patient mentions must remain consistent without exposing real identities.

Practical Code Examples

The following examples demonstrate typical usage patterns for the replace redaction method.

Basic usage with random surrogates:

from openmed import deidentify

text = "Patient John Doe visited on 12/05/2023. Contact: john.doe@example.com"
result = deidentify(text, method="replace", lang="en")
print(result.deidentified_text)

# Output: Patient "Evelyn Owens" visited on 12/05/2023. Contact: "carlos.silva@example.org"

Deterministic replacement ensures reproducibility:


# Same entity always gets the same surrogate

result1 = deidentify(text, method="replace", lang="en", consistent=True, seed=42)
result2 = deidentify(text, method="replace", lang="en", consistent=True, seed=42)
assert result1.deidentified_text == result2.deidentified_text

Multilingual example with Portuguese locale:

result_pt = deidentify(text, method="replace", lang="pt")
print(result_pt.deidentified_text)

# Output: Paciente "Mariana Costa" visitou em 12/05/2023. Contato: "mariana.costa@exemplo.com"

The end-to-end example in examples/privacy_filter_unified.py (lines 86-89) demonstrates production usage with consistent=True and explicit seeding.

Testing and Validation

The test suite in tests/unit/test_pii.py (lines 358-585) comprehensively validates the replace strategy. The test_deidentify_replace_method case verifies that the deidentify function returns a DeidentificationResult containing Faker-generated values rather than original PII. For deterministic behavior, test_pt_replace_seeded_is_repeatable confirms that Portuguese text processing with fixed seeds yields identical outputs across multiple invocations. The internal helper _redact_entity receives direct validation through test_redact_replace, ensuring individual entities transform correctly.

Summary

  • OpenMed's replace strategy generates realistic, locale-aware surrogates using the Faker library, preserving text utility for NLP tasks while removing sensitive PII.
  • Deterministic operation is achieved through the consistent flag and seed parameter, which hash original entity text to ensure repeatable replacements across runs.
  • Core implementation resides in openmed/core/pii.py, specifically around line 930, handling entity detection, surrogate generation, and offset management.
  • Multilingual support automatically adapts surrogate generation to the target language specified via the lang parameter.
  • Comprehensive testing in tests/unit/test_pii.py validates both random and deterministic replacement behaviors.

Frequently Asked Questions

What is the difference between replace and mask redaction in OpenMed?

The mask strategy replaces PII entities with fixed-length asterisks or blocking characters, which destroys contextual information and grammatical structure. The replace strategy substitutes entities with realistic fake data—such as replacing "John Doe" with "Evelyn Owens"—maintaining the text's linguistic patterns and readability. This makes replace superior for training clinical NLP models where realistic sentence structure matters.

How do I ensure the same PII entity always gets the same surrogate across multiple documents?

Pass consistent=True along with a fixed integer seed when calling deidentify(). According to the implementation in openmed/core/pii.py, this hashes the original entity text and combines it with your seed to initialize the Faker generator. Consequently, identical strings like "John Doe" will consistently map to the same surrogate across different documents and processing runs.

Does OpenMed support languages other than English for surrogate generation?

Yes. OpenMed's replace strategy uses the Faker library with locale-specific generators. When you specify lang="pt" for Portuguese, lang="de" for German, or other supported language codes, the system generates culturally appropriate names, addresses, and phone numbers. This is validated in the test suite through cases like test_pt_replace_seeded_is_repeatable.

Can I reverse the replace operation to recover original PII?

The DeidentificationResult object includes an optional mapping for reversible redaction, though this feature depends on your specific implementation and whether you preserve the mapping externally. For strict compliance with HIPAA Safe Harbor provisions, you should treat the de-identified text as permanent and store the original-to-surrogate mapping securely if reversibility is required for your use case.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →