# Implementing PII De-identification with Replace Redaction in OpenMed

> Implement PII de-identification in OpenMed using replace redaction. Generate realistic, locale-aware surrogates with Faker for reproducible clinical data de-identification.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: tutorial
- Published: 2026-06-12

---

**OpenMed's replace redaction strategy substitutes detected PII entities with realistic, locale-aware surrogates generated by the Faker library, supporting deterministic output through seeded hashing for reproducible de-identification across clinical datasets.**

OpenMed provides a flexible pipeline for PII de-identification in clinical text, supporting multiple redaction strategies including **mask**, **remove**, **replace**, **hash**, and **shift_dates**. The replace method offers the most sophisticated approach by generating fake but realistic substitutes for sensitive entities like names, emails, and phone numbers. This article examines how to implement PII de-identification with replace redaction using the OpenMed source code, covering deterministic surrogate generation and multilingual support.

## Understanding the Replace Redaction Strategy

The replace method stands among five available strategies in OpenMed. While **mask** obscures text with asterisks and **remove** deletes entities entirely, **replace** maintains text realism by injecting contextually appropriate fake data. This approach preserves the grammatical structure and semantic patterns of clinical narratives, making the de-identified text suitable for downstream NLP tasks and research datasets.

### Locale-Aware Surrogate Generation

OpenMed leverages the **Faker** library to generate surrogates that match the linguistic context of the source document. When processing Portuguese clinical notes with `lang="pt"`, the system produces Portuguese names and addresses rather than English equivalents. This locale awareness ensures that de-identified datasets maintain the cultural and linguistic characteristics necessary for training language-specific models.

## Implementation Details in OpenMed

### Entry Point and Core Logic

The de-identification pipeline centers on the `deidentify` function in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py). This function orchestrates NER model inference—using either Torch or MLX implementations—and delegates to specific redaction branches based on the `method` parameter. The function returns a `DeidentificationResult` object containing the processed text, surrogate mappings, and entity metadata.

### The Replace Branch Implementation

Around **line 930** in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py), the replace logic resides within an `elif method == "replace":` branch. For each detected `PIIEntity`, the code determines the appropriate Faker locale from the `lang` argument and generates type-specific surrogates by calling methods like `fake.name()`, `fake.email()`, or `fake.phone_number()`. Because surrogate lengths may differ from original entities, the algorithm updates the offset of subsequent entities to maintain correct character positions throughout the document.

### Deterministic Replacement with Consistent Seeding

For research reproducibility, OpenMed supports deterministic surrogate generation through the `consistent` parameter. When `consistent=True` is passed alongside a `seed` value, the code hashes the original entity text and merges it with the user-provided seed. This hash initializes the Faker generator, ensuring that identical entities always map to identical surrogates across multiple runs. This implementation is crucial for longitudinal studies where patient mentions must remain consistent without exposing real identities.

## Practical Code Examples

The following examples demonstrate typical usage patterns for the replace redaction method.

Basic usage with random surrogates:

```python
from openmed import deidentify

text = "Patient John Doe visited on 12/05/2023. Contact: john.doe@example.com"
result = deidentify(text, method="replace", lang="en")
print(result.deidentified_text)

# Output: Patient "Evelyn Owens" visited on 12/05/2023. Contact: "carlos.silva@example.org"

```

Deterministic replacement ensures reproducibility:

```python

# Same entity always gets the same surrogate

result1 = deidentify(text, method="replace", lang="en", consistent=True, seed=42)
result2 = deidentify(text, method="replace", lang="en", consistent=True, seed=42)
assert result1.deidentified_text == result2.deidentified_text

```

Multilingual example with Portuguese locale:

```python
result_pt = deidentify(text, method="replace", lang="pt")
print(result_pt.deidentified_text)

# Output: Paciente "Mariana Costa" visitou em 12/05/2023. Contato: "mariana.costa@exemplo.com"

```

The end-to-end example in [`examples/privacy_filter_unified.py`](https://github.com/maziyarpanahi/openmed/blob/main/examples/privacy_filter_unified.py) (lines 86-89) demonstrates production usage with `consistent=True` and explicit seeding.

## Testing and Validation

The test suite in [`tests/unit/test_pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_pii.py) (lines 358-585) comprehensively validates the replace strategy. The `test_deidentify_replace_method` case verifies that the `deidentify` function returns a `DeidentificationResult` containing Faker-generated values rather than original PII. For deterministic behavior, `test_pt_replace_seeded_is_repeatable` confirms that Portuguese text processing with fixed seeds yields identical outputs across multiple invocations. The internal helper `_redact_entity` receives direct validation through `test_redact_replace`, ensuring individual entities transform correctly.

## Summary

- **OpenMed's replace strategy** generates realistic, locale-aware surrogates using the Faker library, preserving text utility for NLP tasks while removing sensitive PII.
- **Deterministic operation** is achieved through the `consistent` flag and `seed` parameter, which hash original entity text to ensure repeatable replacements across runs.
- **Core implementation** resides in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py), specifically around line 930, handling entity detection, surrogate generation, and offset management.
- **Multilingual support** automatically adapts surrogate generation to the target language specified via the `lang` parameter.
- **Comprehensive testing** in [`tests/unit/test_pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_pii.py) validates both random and deterministic replacement behaviors.

## Frequently Asked Questions

### What is the difference between replace and mask redaction in OpenMed?

The **mask** strategy replaces PII entities with fixed-length asterisks or blocking characters, which destroys contextual information and grammatical structure. The **replace** strategy substitutes entities with realistic fake data—such as replacing "John Doe" with "Evelyn Owens"—maintaining the text's linguistic patterns and readability. This makes replace superior for training clinical NLP models where realistic sentence structure matters.

### How do I ensure the same PII entity always gets the same surrogate across multiple documents?

Pass `consistent=True` along with a fixed integer `seed` when calling `deidentify()`. According to the implementation in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py), this hashes the original entity text and combines it with your seed to initialize the Faker generator. Consequently, identical strings like "John Doe" will consistently map to the same surrogate across different documents and processing runs.

### Does OpenMed support languages other than English for surrogate generation?

Yes. OpenMed's replace strategy uses the Faker library with locale-specific generators. When you specify `lang="pt"` for Portuguese, `lang="de"` for German, or other supported language codes, the system generates culturally appropriate names, addresses, and phone numbers. This is validated in the test suite through cases like `test_pt_replace_seeded_is_repeatable`.

### Can I reverse the replace operation to recover original PII?

The `DeidentificationResult` object includes an optional mapping for reversible redaction, though this feature depends on your specific implementation and whether you preserve the mapping externally. For strict compliance with HIPAA Safe Harbor provisions, you should treat the de-identified text as permanent and store the original-to-surrogate mapping securely if reversibility is required for your use case.