OpenMed Security Best Practices: Deterministic PII Anonymization and Safety Sweeps
OpenMed enforces privacy through deterministic, rule-based post-processing that removes or replaces PII using regex-based safety sweeps, locale-aware anonymization with consistent seeding, and canonical label normalization.
OpenMed is a privacy-preserving medical-text processing library developed by maziyarpanahi. The codebase implements a defense-in-depth security model that combines machine learning detection with deterministic rule-based validation. Understanding these OpenMed security best practices ensures that sensitive patient data never leaks through your NLP pipelines.
Architectural Safeguards in OpenMed
The security architecture revolves around four core components that enforce deterministic PII handling:
| Component | Responsibility | Key File |
|---|---|---|
safety_sweep |
Adds deterministic regex-based PII spans after ML detection | openmed/core/safety_sweep.py |
Anonymizer |
Generates locale-aware fake surrogates with optional deterministic mode | openmed/core/anonymizer/engine.py |
PII Entity Merger |
Normalizes overlapping detections and validates context words | openmed/core/pii_entity_merger.py |
Labels |
Centralizes canonical label catalogue to avoid label sprawl | openmed/core/labels.py |
Deterministic Post-Detector Safety Sweeps
The safety_sweep function in openmed/core/safety_sweep.py runs after any ML model emits spans. Existing spans always win priority, preventing accidental overwrite of higher-confidence detections. Candidates are generated from language-specific regex patterns (PII_PATTERNS) and sorted by confidence, priority, and span size in _collect_candidates (lines 16-23).
Overlap checks in _overlaps (lines 40-47) guarantee that newly added spans never collide with already-selected spans. This ensures deterministic coverage of known PII patterns regardless of ML model variance.
Pattern Validation and Context Awareness
Each PIIPattern may provide a custom validator (_validated) executed before candidate acceptance (lines 59-66 in safety_sweep.py). The _confidence method (lines 68-77) boosts confidence when surrounding context words (e.g., "ssn", "social security") are present.
Best practice: Define strict validators for patterns that could match benign substrings, and supply appropriate context_words to reduce false positives.
Locale-Aware Anonymization with Deterministic Seeding
The Anonymizer class in openmed/core/anonymizer/engine.py builds a Faker instance per locale via _build_faker (lines 101-125) and caches it for reuse. When consistent=True, the _derive_seed method (lines 136-142) seeds the Faker with a hash derived from the canonical label and original value. This guarantees the same surrogate generates for repeated occurrences, eliminating cross-record linking attacks.
Best practice: Enable deterministic mode for document-wide consistency; optionally set a fixed seed parameter if you need reproducibility across runs.
Canonical Label Normalization
Labels normalize to canonical forms (e.g., "social_security_number" → SSN) via normalize_label (see usage at line 72 in engine.py). Centralizing label definitions in openmed/core/labels.py reduces attack surface from misspelled or duplicate label names.
Best practice: Always use the library's normalize_label before downstream processing to avoid accidentally exposing raw label strings.
Secure Integration Workflow
Follow this checklist when embedding OpenMed in production:
- Input sanitisation – Ensure incoming text is UTF-8 and strip unexpected binary payloads.
- Run ML detector – Produce spans in the format expected by
safety_sweep. - Invoke
safety_sweep– Pass the original text and ML spans; keep only the returned non-overlapping spans. - Configure
Anonymizer– Enableconsistent=True; optionally set a fixedseedfor repeatable results. - Locale handling – Pass the correct
langcode and verify the Faker locale resolves (fallbacks log at lines 112-119 ofengine.py). - Never log raw PII – Log only sanitized spans or their metadata (
_candidate_metadataat lines 80-89). - Secure deployment – Run the library in an isolated container without network-exposed credentials; OpenMed itself holds no secrets.
- Audit pattern set – Regularly review
PII_PATTERNSviapii_i18n.get_patterns_for_languagerather than editing core regexes directly.
Implementation Examples
Running Safety Sweeps on ML Detections
from openmed.core.safety_sweep import safety_sweep
text = "Patient John Doe, SSN 123-45-6789, visited on 2024-03-01."
ml_spans = [] # Assume empty for illustration
# Add deterministic regex-based PII spans
secure_spans = safety_sweep(text, ml_spans, lang="en")
The safety_sweep function guarantees no span overlaps existing detections and validates each addition against its pattern via _validated (lines 59-66).
Configuring Deterministic Anonymization
from openmed.core.anonymizer.engine import Anonymizer
anon = Anonymizer(lang="en", consistent=True, seed=42)
# First occurrence generates surrogate
fake_ssn = anon.surrogate("123-45-6789", "social_security_number")
print(fake_ssn) # e.g., "876-54-3210"
# Re-use returns identical surrogate
assert anon.surrogate("123-45-6789", "social_security_number") == fake_ssn
Determinism is achieved via _derive_seed (lines 136-142) using faker.seed_instance.
Building a Complete Secure Pipeline
from openmed.core.safety_sweep import safety_sweep
from openmed.core.anonymizer.engine import Anonymizer
def secure_process(text: str, ml_spans: list):
# 1. Add deterministic PII spans
spans = safety_sweep(text, ml_spans, lang="en")
# 2. Replace with consistent surrogates
anon = Anonymizer(lang="en", consistent=True, seed=2023)
for span in spans:
fake = anon.surrogate(span.text, span.label)
text = text[:span.start] + fake + text[span.end:]
return text
This pattern ensures no raw PII leaves the process and surrogates remain locale-aware and reproducible.
Summary
- Always execute the safety sweep after ML detection to guarantee coverage of known PII patterns via deterministic regex validation.
- Enable deterministic anonymization (
consistent=True) with fixed seeds to prevent cross-record linkage attacks while maintaining consistency. - Use canonical labels from
openmed/core/labels.pyand normalize all inputs to avoid label sprawl vulnerabilities. - Audit patterns responsibly by extending
pii_i18n.get_patterns_for_languagerather than modifying core regexes. - Never log raw PII; rely on
_candidate_metadatafor debugging and audit trails.
Frequently Asked Questions
How does OpenMed prevent ML detection failures from exposing PII?
OpenMed implements a deterministic post-detector sweep in openmed/core/safety_sweep.py that runs after any ML model. The _collect_candidates method (lines 16-23) generates spans from language-specific regex patterns, while _overlaps (lines 40-47) ensures these never overwrite existing high-confidence detections. This guarantees that known PII patterns are always redacted even if the ML model misses them.
What is the difference between consistent and non-consistent anonymization modes?
When consistent=True, the Anonymizer uses _derive_seed (lines 136-142 in engine.py) to hash the canonical label and original value, ensuring identical inputs always produce identical surrogates. This prevents cross-record linkage attacks while maintaining referential integrity within documents. Non-consistent mode generates random surrogates for each occurrence, suitable when referential consistency is not required.
How should I extend OpenMed to support new PII patterns for my locale?
Add new patterns via pii_i18n.get_patterns_for_language rather than editing PII_PATTERNS directly. Each pattern should include a custom validator in _validated (lines 59-66 of safety_sweep.py) and appropriate context_words to reduce false positives. This approach maintains the integrity of the core library while allowing locale-specific extensions.
Is it safe to log the output of OpenMed's processing for debugging?
Only log metadata provided by _candidate_metadata (lines 80-89 in engine.py) or the sanitized spans themselves. Never log raw input text or raw PII values. The library is designed to ensure that once processed through the safety sweep and anonymizer, no sensitive data remains in the output, making it safe for downstream logging and analytics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →