How to Handle Multilingual PII Detection Across 12 Languages in OpenMed
OpenMed provides a plug-and-play multilingual PII detection layer that automatically selects language-specific regex patterns and models based on ISO-639-1 language codes, merging them with a universal fallback set to identify personally identifiable information across diverse text formats.
OpenMed is an open-source medical NLP framework that includes robust multilingual PII detection capabilities for clinical text processing. The system supports 12+ languages through a modular architecture that combines language-specific regular expressions with transformer-based models. This implementation lives primarily in the openmed/core/pii_i18n.py and openmed/core/pii_entity_merger.py modules, providing a seamless API through the PrivacyFilter class.
Core Architecture Components
The multilingual system relies on a layered architecture that separates language-agnostic patterns from localized ones.
Pattern Definitions and Dataclasses
The PIIPattern dataclass, defined in openmed/core/pii_entity_merger.py, encapsulates detection logic with three critical fields: a compiled regex, the entity type (e.g., date, phone_number), and optional language metadata. This structure standardizes how regular expressions are stored and executed across the codebase.
Language-Specific Pattern Registry
The LANGUAGE_PII_PATTERNS dictionary in openmed/core/pii_i18n.py maps ISO-639-1 codes (e.g., "fr", "de", "ja") to lists of PIIPattern objects tailored for regional formatting conventions. This sits alongside PII_PATTERNS, a global list of English-centric patterns that serve as a fallback when language-specific coverage is incomplete.
Model Configuration
The DEFAULT_PII_MODELS dictionary, also in openmed/core/pii_i18n.py, associates language codes with specific transformer model identifiers. For example, French text triggers the OpenMed/OpenMed-PII-SuperClinical-French-44M-v1 model, ensuring tokenization and NER heads match the target language's morphology.
The Detection Pipeline
When calling privacy.analyze_text(text, language="fr"), the PrivacyFilter class in openmed/mlx/models/privacy_filter.py executes a five-stage pipeline:
- Model Resolution – The system looks up
DEFAULT_PII_MODELS[lang]to load the appropriate language-specific transformer. - Pattern Merging –
LANGUAGE_PII_PATTERNS[lang]is concatenated withPII_PATTERNS, ensuring language-specific regexes are evaluated first. - Regex Scanning – Each
PIIPattern.regexis applied to raw text using pre-compiled patterns to minimize overhead. - Entity Merging – Raw matches pass through
merge_entities_with_semantic_unitsinopenmed/core/pii_entity_merger.py, which normalizes dates, deduplicates overlapping spans, and resolves conflicts. - Schema Normalization – Results return as a language-agnostic list containing
entity_type,text,start,end, andlanguagefields.
Handling Language-Specific Edge Cases
The system addresses regional formatting variations through specialized patterns in LANGUAGE_PII_PATTERNS:
- Date Formats – German dates (
dd.mm.yyyy) and Japanese dates (yyyy-mm-dd) use distinctPIIPatternentries. The merger normalizes these to ISO-8601 regardless of input format. - Phone Number Separators – Patterns include optional spaces, dashes, and parentheses, while the merger strips non-numeric characters to produce canonical numbers.
- National ID Validation – Language-specific regexes capture full ID structures, with optional checksum validation (e.g., for Indian Aadhaar numbers) during the merging phase.
- Overlapping Entities – When a date appears inside a larger address span, the
is_more_specificranking inmerge_entities_with_semantic_unitspreserves the narrower detection while annotating the broader context.
Practical Implementation Examples
Initialize a Language-Specific Filter
Load the French PII detection pipeline by specifying the ISO-639-1 code:
from openmed.mlx.models.privacy_filter import PrivacyFilter
# Automatically resolves model from DEFAULT_PII_MODELS
privacy = PrivacyFilter(language="fr")
Analyze French Clinical Text
Process text containing region-specific identifiers:
clinical_text = """
Patient né le 12/03/1975, adresse 12 rue de la Paix, 75002 Paris.
Numéro de sécurité sociale : 1 84 12 75 123 456 78.
"""
entities = privacy.analyze_text(clinical_text)
for entity in entities:
print(f"{entity['entity_type']}: {entity['text']} ({entity['language']})")
Expected output:
[
{"entity_type": "date", "text": "12/03/1975", "start": 12, "end": 22, "language": "fr"},
{"entity_type": "street_address", "text": "12 rue de la Paix", "start": 28, "end": 45, "language": "fr"},
{"entity_type": "national_id", "text": "1 84 12 75 123 456 78", "start": 73, "end": 96, "language": "fr"}
]
Extend Support for a New Language
Add Swahili support by extending the configuration dictionaries:
# In openmed/core/pii_i18n.py:
# 1. Add patterns to LANGUAGE_PII_PATTERNS
LANGUAGE_PII_PATTERNS["sw"] = [
PIIPattern(
regex=re.compile(r'\b\d{2}/\d{2}/\d{4}\b'),
entity_type="date",
language="sw"
)
]
# 2. Register model in DEFAULT_PII_MODELS
DEFAULT_PII_MODELS["sw"] = "OpenMed/OpenMed-PII-Swahili-44M-v1"
After modification, validate the integration using the regression test suite:
pytest tests/unit/test_pii_i18n.py -v
Summary
- Multilingual PII detection in OpenMed relies on the
LANGUAGE_PII_PATTERNSdictionary inopenmed/core/pii_i18n.pyto store language-specific regex configurations. - The
PrivacyFilterclass automatically selects appropriate models and patterns based on ISO-639-1 language codes, falling back to English-centric patterns when necessary. - Entity merging occurs through
merge_entities_with_semantic_unitsinopenmed/core/pii_entity_merger.py, which handles overlapping detections and normalizes regional formatting variations. - Adding support for new languages requires only extending
LANGUAGE_PII_PATTERNSandDEFAULT_PII_MODELS, with validation viatests/unit/test_pii_i18n.py.
Frequently Asked Questions
How does OpenMed handle unsupported languages?
If a language code is absent from DEFAULT_PII_MODELS, the system falls back to the generic multilingual model defined under the "en" key while still applying any available language-specific regexes from LANGUAGE_PII_PATTERNS. This ensures graceful degradation when processing low-resource languages not yet assigned dedicated transformer models.
What is the difference between PII_PATTERNS and LANGUAGE_PII_PATTERNS?
PII_PATTERNS contains language-agnostic regular expressions that serve as universal fallbacks, while LANGUAGE_PII_PATTERNS is a dictionary mapping ISO-639-1 codes to lists of patterns optimized for regional formats like French dates (jj/mm/aaaa) or German phone numbers. During detection, language-specific patterns are evaluated first, followed by the generic fallback set.
How are overlapping PII entities resolved in multilingual text?
The merge_entities_with_semantic_units function ranks potential overlaps using an is_more_specific heuristic, preserving granular detections (like dates within addresses) while preventing duplicate annotations. It normalizes formatting differences—such as stripping separators from phone numbers—before returning the final entity list.
Can I use custom regex patterns for specific languages?
Yes. You can append PIIPattern instances to any language entry in LANGUAGE_PII_PATTERNS before initializing PrivacyFilter. These custom patterns are compiled once per language and evaluated alongside the built-in patterns, allowing domain-specific extensions (e.g., hospital-specific ID formats) without modifying the core library.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →