# How to Handle Multilingual PII Detection Across 12 Languages in OpenMed

> Learn how OpenMed handles multilingual PII detection across 12 languages using language-specific regex and models with a universal fallback for accurate identification in any text.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**OpenMed provides a plug-and-play multilingual PII detection layer that automatically selects language-specific regex patterns and models based on ISO-639-1 language codes, merging them with a universal fallback set to identify personally identifiable information across diverse text formats.**

OpenMed is an open-source medical NLP framework that includes robust **multilingual PII detection** capabilities for clinical text processing. The system supports 12+ languages through a modular architecture that combines language-specific regular expressions with transformer-based models. This implementation lives primarily in the [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py) and [`openmed/core/pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_entity_merger.py) modules, providing a seamless API through the `PrivacyFilter` class.

## Core Architecture Components

The multilingual system relies on a layered architecture that separates language-agnostic patterns from localized ones.

### Pattern Definitions and Dataclasses

The **`PIIPattern`** dataclass, defined in [`openmed/core/pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_entity_merger.py), encapsulates detection logic with three critical fields: a compiled regex, the entity type (e.g., `date`, `phone_number`), and optional language metadata. This structure standardizes how regular expressions are stored and executed across the codebase.

### Language-Specific Pattern Registry

The **`LANGUAGE_PII_PATTERNS`** dictionary in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py) maps ISO-639-1 codes (e.g., `"fr"`, `"de"`, `"ja"`) to lists of `PIIPattern` objects tailored for regional formatting conventions. This sits alongside **`PII_PATTERNS`**, a global list of English-centric patterns that serve as a fallback when language-specific coverage is incomplete.

### Model Configuration

The **`DEFAULT_PII_MODELS`** dictionary, also in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py), associates language codes with specific transformer model identifiers. For example, French text triggers the `OpenMed/OpenMed-PII-SuperClinical-French-44M-v1` model, ensuring tokenization and NER heads match the target language's morphology.

## The Detection Pipeline

When calling `privacy.analyze_text(text, language="fr")`, the `PrivacyFilter` class in [`openmed/mlx/models/privacy_filter.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/models/privacy_filter.py) executes a five-stage pipeline:

1. **Model Resolution** – The system looks up `DEFAULT_PII_MODELS[lang]` to load the appropriate language-specific transformer.
2. **Pattern Merging** – `LANGUAGE_PII_PATTERNS[lang]` is concatenated with `PII_PATTERNS`, ensuring language-specific regexes are evaluated first.
3. **Regex Scanning** – Each `PIIPattern.regex` is applied to raw text using pre-compiled patterns to minimize overhead.
4. **Entity Merging** – Raw matches pass through `merge_entities_with_semantic_units` in [`openmed/core/pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_entity_merger.py), which normalizes dates, deduplicates overlapping spans, and resolves conflicts.
5. **Schema Normalization** – Results return as a language-agnostic list containing `entity_type`, `text`, `start`, `end`, and `language` fields.

## Handling Language-Specific Edge Cases

The system addresses regional formatting variations through specialized patterns in `LANGUAGE_PII_PATTERNS`:

- **Date Formats** – German dates (`dd.mm.yyyy`) and Japanese dates (`yyyy-mm-dd`) use distinct `PIIPattern` entries. The merger normalizes these to ISO-8601 regardless of input format.
- **Phone Number Separators** – Patterns include optional spaces, dashes, and parentheses, while the merger strips non-numeric characters to produce canonical numbers.
- **National ID Validation** – Language-specific regexes capture full ID structures, with optional checksum validation (e.g., for Indian Aadhaar numbers) during the merging phase.
- **Overlapping Entities** – When a date appears inside a larger address span, the `is_more_specific` ranking in `merge_entities_with_semantic_units` preserves the narrower detection while annotating the broader context.

## Practical Implementation Examples

### Initialize a Language-Specific Filter

Load the French PII detection pipeline by specifying the ISO-639-1 code:

```python
from openmed.mlx.models.privacy_filter import PrivacyFilter

# Automatically resolves model from DEFAULT_PII_MODELS

privacy = PrivacyFilter(language="fr")

```

### Analyze French Clinical Text

Process text containing region-specific identifiers:

```python
clinical_text = """
Patient né le 12/03/1975, adresse 12 rue de la Paix, 75002 Paris.
Numéro de sécurité sociale : 1 84 12 75 123 456 78.
"""

entities = privacy.analyze_text(clinical_text)
for entity in entities:
    print(f"{entity['entity_type']}: {entity['text']} ({entity['language']})")

```

Expected output:

```json
[
  {"entity_type": "date", "text": "12/03/1975", "start": 12, "end": 22, "language": "fr"},
  {"entity_type": "street_address", "text": "12 rue de la Paix", "start": 28, "end": 45, "language": "fr"},
  {"entity_type": "national_id", "text": "1 84 12 75 123 456 78", "start": 73, "end": 96, "language": "fr"}
]

```

### Extend Support for a New Language

Add Swahili support by extending the configuration dictionaries:

```python

# In openmed/core/pii_i18n.py:

# 1. Add patterns to LANGUAGE_PII_PATTERNS

LANGUAGE_PII_PATTERNS["sw"] = [
    PIIPattern(
        regex=re.compile(r'\b\d{2}/\d{2}/\d{4}\b'),
        entity_type="date",
        language="sw"
    )
]

# 2. Register model in DEFAULT_PII_MODELS

DEFAULT_PII_MODELS["sw"] = "OpenMed/OpenMed-PII-Swahili-44M-v1"

```

After modification, validate the integration using the regression test suite:

```bash
pytest tests/unit/test_pii_i18n.py -v

```

## Summary

- **Multilingual PII detection** in OpenMed relies on the `LANGUAGE_PII_PATTERNS` dictionary in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py) to store language-specific regex configurations.
- The **`PrivacyFilter`** class automatically selects appropriate models and patterns based on ISO-639-1 language codes, falling back to English-centric patterns when necessary.
- **Entity merging** occurs through `merge_entities_with_semantic_units` in [`openmed/core/pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_entity_merger.py), which handles overlapping detections and normalizes regional formatting variations.
- Adding support for new languages requires only extending `LANGUAGE_PII_PATTERNS` and `DEFAULT_PII_MODELS`, with validation via [`tests/unit/test_pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_pii_i18n.py).

## Frequently Asked Questions

### How does OpenMed handle unsupported languages?

If a language code is absent from `DEFAULT_PII_MODELS`, the system falls back to the generic multilingual model defined under the `"en"` key while still applying any available language-specific regexes from `LANGUAGE_PII_PATTERNS`. This ensures graceful degradation when processing low-resource languages not yet assigned dedicated transformer models.

### What is the difference between PII_PATTERNS and LANGUAGE_PII_PATTERNS?

`PII_PATTERNS` contains language-agnostic regular expressions that serve as universal fallbacks, while `LANGUAGE_PII_PATTERNS` is a dictionary mapping ISO-639-1 codes to lists of patterns optimized for regional formats like French dates (`jj/mm/aaaa`) or German phone numbers. During detection, language-specific patterns are evaluated first, followed by the generic fallback set.

### How are overlapping PII entities resolved in multilingual text?

The `merge_entities_with_semantic_units` function ranks potential overlaps using an `is_more_specific` heuristic, preserving granular detections (like dates within addresses) while preventing duplicate annotations. It normalizes formatting differences—such as stripping separators from phone numbers—before returning the final entity list.

### Can I use custom regex patterns for specific languages?

Yes. You can append `PIIPattern` instances to any language entry in `LANGUAGE_PII_PATTERNS` before initializing `PrivacyFilter`. These custom patterns are compiled once per language and evaluated alongside the built-in patterns, allowing domain-specific extensions (e.g., hospital-specific ID formats) without modifying the core library.