How to Handle Multilingual PII Detection Across 12 Languages in OpenMed

OpenMed provides a plug-and-play multilingual PII detection layer that automatically selects language-specific regex patterns and models based on ISO-639-1 language codes, merging them with a universal fallback set to identify personally identifiable information across diverse text formats.

OpenMed is an open-source medical NLP framework that includes robust multilingual PII detection capabilities for clinical text processing. The system supports 12+ languages through a modular architecture that combines language-specific regular expressions with transformer-based models. This implementation lives primarily in the openmed/core/pii_i18n.py and openmed/core/pii_entity_merger.py modules, providing a seamless API through the PrivacyFilter class.

Core Architecture Components

The multilingual system relies on a layered architecture that separates language-agnostic patterns from localized ones.

Pattern Definitions and Dataclasses

The PIIPattern dataclass, defined in openmed/core/pii_entity_merger.py, encapsulates detection logic with three critical fields: a compiled regex, the entity type (e.g., date, phone_number), and optional language metadata. This structure standardizes how regular expressions are stored and executed across the codebase.

Language-Specific Pattern Registry

The LANGUAGE_PII_PATTERNS dictionary in openmed/core/pii_i18n.py maps ISO-639-1 codes (e.g., "fr", "de", "ja") to lists of PIIPattern objects tailored for regional formatting conventions. This sits alongside PII_PATTERNS, a global list of English-centric patterns that serve as a fallback when language-specific coverage is incomplete.

Model Configuration

The DEFAULT_PII_MODELS dictionary, also in openmed/core/pii_i18n.py, associates language codes with specific transformer model identifiers. For example, French text triggers the OpenMed/OpenMed-PII-SuperClinical-French-44M-v1 model, ensuring tokenization and NER heads match the target language's morphology.

The Detection Pipeline

When calling privacy.analyze_text(text, language="fr"), the PrivacyFilter class in openmed/mlx/models/privacy_filter.py executes a five-stage pipeline:

  1. Model Resolution – The system looks up DEFAULT_PII_MODELS[lang] to load the appropriate language-specific transformer.
  2. Pattern Merging – LANGUAGE_PII_PATTERNS[lang] is concatenated with PII_PATTERNS, ensuring language-specific regexes are evaluated first.
  3. Regex Scanning – Each PIIPattern.regex is applied to raw text using pre-compiled patterns to minimize overhead.
  4. Entity Merging – Raw matches pass through merge_entities_with_semantic_units in openmed/core/pii_entity_merger.py, which normalizes dates, deduplicates overlapping spans, and resolves conflicts.
  5. Schema Normalization – Results return as a language-agnostic list containing entity_type, text, start, end, and language fields.

Handling Language-Specific Edge Cases

The system addresses regional formatting variations through specialized patterns in LANGUAGE_PII_PATTERNS:

  • Date Formats – German dates (dd.mm.yyyy) and Japanese dates (yyyy-mm-dd) use distinct PIIPattern entries. The merger normalizes these to ISO-8601 regardless of input format.
  • Phone Number Separators – Patterns include optional spaces, dashes, and parentheses, while the merger strips non-numeric characters to produce canonical numbers.
  • National ID Validation – Language-specific regexes capture full ID structures, with optional checksum validation (e.g., for Indian Aadhaar numbers) during the merging phase.
  • Overlapping Entities – When a date appears inside a larger address span, the is_more_specific ranking in merge_entities_with_semantic_units preserves the narrower detection while annotating the broader context.

Practical Implementation Examples

Initialize a Language-Specific Filter

Load the French PII detection pipeline by specifying the ISO-639-1 code:

from openmed.mlx.models.privacy_filter import PrivacyFilter

# Automatically resolves model from DEFAULT_PII_MODELS

privacy = PrivacyFilter(language="fr")

Analyze French Clinical Text

Process text containing region-specific identifiers:

clinical_text = """
Patient né le 12/03/1975, adresse 12 rue de la Paix, 75002 Paris.
Numéro de sécurité sociale : 1 84 12 75 123 456 78.
"""

entities = privacy.analyze_text(clinical_text)
for entity in entities:
    print(f"{entity['entity_type']}: {entity['text']} ({entity['language']})")

Expected output:

[
  {"entity_type": "date", "text": "12/03/1975", "start": 12, "end": 22, "language": "fr"},
  {"entity_type": "street_address", "text": "12 rue de la Paix", "start": 28, "end": 45, "language": "fr"},
  {"entity_type": "national_id", "text": "1 84 12 75 123 456 78", "start": 73, "end": 96, "language": "fr"}
]

Extend Support for a New Language

Add Swahili support by extending the configuration dictionaries:


# In openmed/core/pii_i18n.py:

# 1. Add patterns to LANGUAGE_PII_PATTERNS

LANGUAGE_PII_PATTERNS["sw"] = [
    PIIPattern(
        regex=re.compile(r'\b\d{2}/\d{2}/\d{4}\b'),
        entity_type="date",
        language="sw"
    )
]

# 2. Register model in DEFAULT_PII_MODELS

DEFAULT_PII_MODELS["sw"] = "OpenMed/OpenMed-PII-Swahili-44M-v1"

After modification, validate the integration using the regression test suite:

pytest tests/unit/test_pii_i18n.py -v

Summary

  • Multilingual PII detection in OpenMed relies on the LANGUAGE_PII_PATTERNS dictionary in openmed/core/pii_i18n.py to store language-specific regex configurations.
  • The PrivacyFilter class automatically selects appropriate models and patterns based on ISO-639-1 language codes, falling back to English-centric patterns when necessary.
  • Entity merging occurs through merge_entities_with_semantic_units in openmed/core/pii_entity_merger.py, which handles overlapping detections and normalizes regional formatting variations.
  • Adding support for new languages requires only extending LANGUAGE_PII_PATTERNS and DEFAULT_PII_MODELS, with validation via tests/unit/test_pii_i18n.py.

Frequently Asked Questions

How does OpenMed handle unsupported languages?

If a language code is absent from DEFAULT_PII_MODELS, the system falls back to the generic multilingual model defined under the "en" key while still applying any available language-specific regexes from LANGUAGE_PII_PATTERNS. This ensures graceful degradation when processing low-resource languages not yet assigned dedicated transformer models.

What is the difference between PII_PATTERNS and LANGUAGE_PII_PATTERNS?

PII_PATTERNS contains language-agnostic regular expressions that serve as universal fallbacks, while LANGUAGE_PII_PATTERNS is a dictionary mapping ISO-639-1 codes to lists of patterns optimized for regional formats like French dates (jj/mm/aaaa) or German phone numbers. During detection, language-specific patterns are evaluated first, followed by the generic fallback set.

How are overlapping PII entities resolved in multilingual text?

The merge_entities_with_semantic_units function ranks potential overlaps using an is_more_specific heuristic, preserving granular detections (like dates within addresses) while preventing duplicate annotations. It normalizes formatting differences—such as stripping separators from phone numbers—before returning the final entity list.

Can I use custom regex patterns for specific languages?

Yes. You can append PIIPattern instances to any language entry in LANGUAGE_PII_PATTERNS before initializing PrivacyFilter. These custom patterns are compiled once per language and evaluated alongside the built-in patterns, allowing domain-specific extensions (e.g., hospital-specific ID formats) without modifying the core library.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →