OpenMed Security Best Practices: Deterministic PII Anonymization and Safety Sweeps

OpenMed enforces privacy through deterministic, rule-based post-processing that removes or replaces PII using regex-based safety sweeps, locale-aware anonymization with consistent seeding, and canonical label normalization.

OpenMed is a privacy-preserving medical-text processing library developed by maziyarpanahi. The codebase implements a defense-in-depth security model that combines machine learning detection with deterministic rule-based validation. Understanding these OpenMed security best practices ensures that sensitive patient data never leaks through your NLP pipelines.

Architectural Safeguards in OpenMed

The security architecture revolves around four core components that enforce deterministic PII handling:

Component Responsibility Key File
safety_sweep Adds deterministic regex-based PII spans after ML detection openmed/core/safety_sweep.py
Anonymizer Generates locale-aware fake surrogates with optional deterministic mode openmed/core/anonymizer/engine.py
PII Entity Merger Normalizes overlapping detections and validates context words openmed/core/pii_entity_merger.py
Labels Centralizes canonical label catalogue to avoid label sprawl openmed/core/labels.py

Deterministic Post-Detector Safety Sweeps

The safety_sweep function in openmed/core/safety_sweep.py runs after any ML model emits spans. Existing spans always win priority, preventing accidental overwrite of higher-confidence detections. Candidates are generated from language-specific regex patterns (PII_PATTERNS) and sorted by confidence, priority, and span size in _collect_candidates (lines 16-23).

Overlap checks in _overlaps (lines 40-47) guarantee that newly added spans never collide with already-selected spans. This ensures deterministic coverage of known PII patterns regardless of ML model variance.

Pattern Validation and Context Awareness

Each PIIPattern may provide a custom validator (_validated) executed before candidate acceptance (lines 59-66 in safety_sweep.py). The _confidence method (lines 68-77) boosts confidence when surrounding context words (e.g., "ssn", "social security") are present.

Best practice: Define strict validators for patterns that could match benign substrings, and supply appropriate context_words to reduce false positives.

Locale-Aware Anonymization with Deterministic Seeding

The Anonymizer class in openmed/core/anonymizer/engine.py builds a Faker instance per locale via _build_faker (lines 101-125) and caches it for reuse. When consistent=True, the _derive_seed method (lines 136-142) seeds the Faker with a hash derived from the canonical label and original value. This guarantees the same surrogate generates for repeated occurrences, eliminating cross-record linking attacks.

Best practice: Enable deterministic mode for document-wide consistency; optionally set a fixed seed parameter if you need reproducibility across runs.

Canonical Label Normalization

Labels normalize to canonical forms (e.g., "social_security_number" → SSN) via normalize_label (see usage at line 72 in engine.py). Centralizing label definitions in openmed/core/labels.py reduces attack surface from misspelled or duplicate label names.

Best practice: Always use the library's normalize_label before downstream processing to avoid accidentally exposing raw label strings.

Secure Integration Workflow

Follow this checklist when embedding OpenMed in production:

  • Input sanitisation – Ensure incoming text is UTF-8 and strip unexpected binary payloads.
  • Run ML detector – Produce spans in the format expected by safety_sweep.
  • Invoke safety_sweep – Pass the original text and ML spans; keep only the returned non-overlapping spans.
  • Configure Anonymizer – Enable consistent=True; optionally set a fixed seed for repeatable results.
  • Locale handling – Pass the correct lang code and verify the Faker locale resolves (fallbacks log at lines 112-119 of engine.py).
  • Never log raw PII – Log only sanitized spans or their metadata (_candidate_metadata at lines 80-89).
  • Secure deployment – Run the library in an isolated container without network-exposed credentials; OpenMed itself holds no secrets.
  • Audit pattern set – Regularly review PII_PATTERNS via pii_i18n.get_patterns_for_language rather than editing core regexes directly.

Implementation Examples

Running Safety Sweeps on ML Detections

from openmed.core.safety_sweep import safety_sweep

text = "Patient John Doe, SSN 123-45-6789, visited on 2024-03-01."
ml_spans = []  # Assume empty for illustration

# Add deterministic regex-based PII spans

secure_spans = safety_sweep(text, ml_spans, lang="en")

The safety_sweep function guarantees no span overlaps existing detections and validates each addition against its pattern via _validated (lines 59-66).

Configuring Deterministic Anonymization

from openmed.core.anonymizer.engine import Anonymizer

anon = Anonymizer(lang="en", consistent=True, seed=42)

# First occurrence generates surrogate

fake_ssn = anon.surrogate("123-45-6789", "social_security_number")
print(fake_ssn)  # e.g., "876-54-3210"

# Re-use returns identical surrogate

assert anon.surrogate("123-45-6789", "social_security_number") == fake_ssn

Determinism is achieved via _derive_seed (lines 136-142) using faker.seed_instance.

Building a Complete Secure Pipeline

from openmed.core.safety_sweep import safety_sweep
from openmed.core.anonymizer.engine import Anonymizer

def secure_process(text: str, ml_spans: list):
    # 1. Add deterministic PII spans

    spans = safety_sweep(text, ml_spans, lang="en")
    
    # 2. Replace with consistent surrogates

    anon = Anonymizer(lang="en", consistent=True, seed=2023)
    for span in spans:
        fake = anon.surrogate(span.text, span.label)
        text = text[:span.start] + fake + text[span.end:]
    
    return text

This pattern ensures no raw PII leaves the process and surrogates remain locale-aware and reproducible.

Summary

  • Always execute the safety sweep after ML detection to guarantee coverage of known PII patterns via deterministic regex validation.
  • Enable deterministic anonymization (consistent=True) with fixed seeds to prevent cross-record linkage attacks while maintaining consistency.
  • Use canonical labels from openmed/core/labels.py and normalize all inputs to avoid label sprawl vulnerabilities.
  • Audit patterns responsibly by extending pii_i18n.get_patterns_for_language rather than modifying core regexes.
  • Never log raw PII; rely on _candidate_metadata for debugging and audit trails.

Frequently Asked Questions

How does OpenMed prevent ML detection failures from exposing PII?

OpenMed implements a deterministic post-detector sweep in openmed/core/safety_sweep.py that runs after any ML model. The _collect_candidates method (lines 16-23) generates spans from language-specific regex patterns, while _overlaps (lines 40-47) ensures these never overwrite existing high-confidence detections. This guarantees that known PII patterns are always redacted even if the ML model misses them.

What is the difference between consistent and non-consistent anonymization modes?

When consistent=True, the Anonymizer uses _derive_seed (lines 136-142 in engine.py) to hash the canonical label and original value, ensuring identical inputs always produce identical surrogates. This prevents cross-record linkage attacks while maintaining referential integrity within documents. Non-consistent mode generates random surrogates for each occurrence, suitable when referential consistency is not required.

How should I extend OpenMed to support new PII patterns for my locale?

Add new patterns via pii_i18n.get_patterns_for_language rather than editing PII_PATTERNS directly. Each pattern should include a custom validator in _validated (lines 59-66 of safety_sweep.py) and appropriate context_words to reduce false positives. This approach maintains the integrity of the core library while allowing locale-specific extensions.

Is it safe to log the output of OpenMed's processing for debugging?

Only log metadata provided by _candidate_metadata (lines 80-89 in engine.py) or the sanitized spans themselves. Never log raw input text or raw PII values. The library is designed to ensure that once processed through the safety sweep and anonymizer, no sensitive data remains in the output, making it safe for downstream logging and analytics.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →