# OpenMed Security Best Practices: Deterministic PII Anonymization and Safety Sweeps

> Discover OpenMed security best practices. Learn how deterministic PII anonymization and safety sweeps ensure data privacy with regex, locale-aware masking, and consistent seeding.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: best-practices
- Published: 2026-06-13

---

**OpenMed enforces privacy through deterministic, rule-based post-processing that removes or replaces PII using regex-based safety sweeps, locale-aware anonymization with consistent seeding, and canonical label normalization.**

OpenMed is a privacy-preserving medical-text processing library developed by maziyarpanahi. The codebase implements a defense-in-depth security model that combines machine learning detection with deterministic rule-based validation. Understanding these OpenMed security best practices ensures that sensitive patient data never leaks through your NLP pipelines.

## Architectural Safeguards in OpenMed

The security architecture revolves around four core components that enforce deterministic PII handling:

| Component | Responsibility | Key File |
|-----------|---------------|----------|
| `safety_sweep` | Adds deterministic regex-based PII spans after ML detection | [`openmed/core/safety_sweep.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/safety_sweep.py) |
| `Anonymizer` | Generates locale-aware fake surrogates with optional deterministic mode | [`openmed/core/anonymizer/engine.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/anonymizer/engine.py) |
| `PII Entity Merger` | Normalizes overlapping detections and validates context words | [`openmed/core/pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_entity_merger.py) |
| `Labels` | Centralizes canonical label catalogue to avoid label sprawl | [`openmed/core/labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/labels.py) |

### Deterministic Post-Detector Safety Sweeps

The `safety_sweep` function in [`openmed/core/safety_sweep.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/safety_sweep.py) runs **after** any ML model emits spans. Existing spans always win priority, preventing accidental overwrite of higher-confidence detections. Candidates are generated from language-specific regex patterns (`PII_PATTERNS`) and sorted by confidence, priority, and span size in `_collect_candidates` (lines 16-23).

Overlap checks in `_overlaps` (lines 40-47) guarantee that newly added spans never collide with already-selected spans. This ensures deterministic coverage of known PII patterns regardless of ML model variance.

### Pattern Validation and Context Awareness

Each `PIIPattern` may provide a custom validator (`_validated`) executed before candidate acceptance (lines 59-66 in [`safety_sweep.py`](https://github.com/maziyarpanahi/openmed/blob/main/safety_sweep.py)). The `_confidence` method (lines 68-77) boosts confidence when surrounding context words (e.g., "ssn", "social security") are present.

**Best practice:** Define strict validators for patterns that could match benign substrings, and supply appropriate `context_words` to reduce false positives.

### Locale-Aware Anonymization with Deterministic Seeding

The `Anonymizer` class in [`openmed/core/anonymizer/engine.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/anonymizer/engine.py) builds a **Faker** instance per locale via `_build_faker` (lines 101-125) and caches it for reuse. When `consistent=True`, the `_derive_seed` method (lines 136-142) seeds the Faker with a hash derived from the canonical label and original value. This guarantees the same surrogate generates for repeated occurrences, eliminating cross-record linking attacks.

**Best practice:** Enable deterministic mode for document-wide consistency; optionally set a fixed `seed` parameter if you need reproducibility across runs.

### Canonical Label Normalization

Labels normalize to canonical forms (e.g., `"social_security_number"` → `SSN`) via `normalize_label` (see usage at line 72 in [`engine.py`](https://github.com/maziyarpanahi/openmed/blob/main/engine.py)). Centralizing label definitions in [`openmed/core/labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/labels.py) reduces attack surface from misspelled or duplicate label names.

**Best practice:** Always use the library's `normalize_label` before downstream processing to avoid accidentally exposing raw label strings.

## Secure Integration Workflow

Follow this checklist when embedding OpenMed in production:

- **Input sanitisation** – Ensure incoming text is UTF-8 and strip unexpected binary payloads.
- **Run ML detector** – Produce spans in the format expected by `safety_sweep`.
- **Invoke `safety_sweep`** – Pass the original text and ML spans; keep only the returned non-overlapping spans.
- **Configure `Anonymizer`** – Enable `consistent=True`; optionally set a fixed `seed` for repeatable results.
- **Locale handling** – Pass the correct `lang` code and verify the Faker locale resolves (fallbacks log at lines 112-119 of [`engine.py`](https://github.com/maziyarpanahi/openmed/blob/main/engine.py)).
- **Never log raw PII** – Log only sanitized spans or their metadata (`_candidate_metadata` at lines 80-89).
- **Secure deployment** – Run the library in an isolated container without network-exposed credentials; OpenMed itself holds no secrets.
- **Audit pattern set** – Regularly review `PII_PATTERNS` via `pii_i18n.get_patterns_for_language` rather than editing core regexes directly.

## Implementation Examples

### Running Safety Sweeps on ML Detections

```python
from openmed.core.safety_sweep import safety_sweep

text = "Patient John Doe, SSN 123-45-6789, visited on 2024-03-01."
ml_spans = []  # Assume empty for illustration

# Add deterministic regex-based PII spans

secure_spans = safety_sweep(text, ml_spans, lang="en")

```

The `safety_sweep` function guarantees no span overlaps existing detections and validates each addition against its pattern via `_validated` (lines 59-66).

### Configuring Deterministic Anonymization

```python
from openmed.core.anonymizer.engine import Anonymizer

anon = Anonymizer(lang="en", consistent=True, seed=42)

# First occurrence generates surrogate

fake_ssn = anon.surrogate("123-45-6789", "social_security_number")
print(fake_ssn)  # e.g., "876-54-3210"

# Re-use returns identical surrogate

assert anon.surrogate("123-45-6789", "social_security_number") == fake_ssn

```

Determinism is achieved via `_derive_seed` (lines 136-142) using `faker.seed_instance`.

### Building a Complete Secure Pipeline

```python
from openmed.core.safety_sweep import safety_sweep
from openmed.core.anonymizer.engine import Anonymizer

def secure_process(text: str, ml_spans: list):
    # 1. Add deterministic PII spans

    spans = safety_sweep(text, ml_spans, lang="en")
    
    # 2. Replace with consistent surrogates

    anon = Anonymizer(lang="en", consistent=True, seed=2023)
    for span in spans:
        fake = anon.surrogate(span.text, span.label)
        text = text[:span.start] + fake + text[span.end:]
    
    return text

```

This pattern ensures **no raw PII leaves the process** and surrogates remain locale-aware and reproducible.

## Summary

- **Always execute the safety sweep** after ML detection to guarantee coverage of known PII patterns via deterministic regex validation.
- **Enable deterministic anonymization** (`consistent=True`) with fixed seeds to prevent cross-record linkage attacks while maintaining consistency.
- **Use canonical labels** from [`openmed/core/labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/labels.py) and normalize all inputs to avoid label sprawl vulnerabilities.
- **Audit patterns responsibly** by extending `pii_i18n.get_patterns_for_language` rather than modifying core regexes.
- **Never log raw PII**; rely on `_candidate_metadata` for debugging and audit trails.

## Frequently Asked Questions

### How does OpenMed prevent ML detection failures from exposing PII?

OpenMed implements a **deterministic post-detector sweep** in [`openmed/core/safety_sweep.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/safety_sweep.py) that runs after any ML model. The `_collect_candidates` method (lines 16-23) generates spans from language-specific regex patterns, while `_overlaps` (lines 40-47) ensures these never overwrite existing high-confidence detections. This guarantees that known PII patterns are always redacted even if the ML model misses them.

### What is the difference between consistent and non-consistent anonymization modes?

When `consistent=True`, the `Anonymizer` uses `_derive_seed` (lines 136-142 in [`engine.py`](https://github.com/maziyarpanahi/openmed/blob/main/engine.py)) to hash the canonical label and original value, ensuring identical inputs always produce identical surrogates. This prevents cross-record linkage attacks while maintaining referential integrity within documents. Non-consistent mode generates random surrogates for each occurrence, suitable when referential consistency is not required.

### How should I extend OpenMed to support new PII patterns for my locale?

Add new patterns via `pii_i18n.get_patterns_for_language` rather than editing `PII_PATTERNS` directly. Each pattern should include a custom validator in `_validated` (lines 59-66 of [`safety_sweep.py`](https://github.com/maziyarpanahi/openmed/blob/main/safety_sweep.py)) and appropriate `context_words` to reduce false positives. This approach maintains the integrity of the core library while allowing locale-specific extensions.

### Is it safe to log the output of OpenMed's processing for debugging?

Only log metadata provided by `_candidate_metadata` (lines 80-89 in [`engine.py`](https://github.com/maziyarpanahi/openmed/blob/main/engine.py)) or the sanitized spans themselves. **Never log raw input text** or raw PII values. The library is designed to ensure that once processed through the safety sweep and anonymizer, no sensitive data remains in the output, making it safe for downstream logging and analytics.