# How to Customize Clinical ID Providers (CPF, CNPJ, BSN, etc.) in OpenMed

> Learn to customize clinical ID providers like CPF, CNPJ, and BSN in OpenMed by modifying pii_i18n.py and defining custom regex patterns and validator functions for accurate patient identification.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**Customize clinical ID providers in OpenMed by implementing validator functions in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py) and linking them to `PIIPattern` objects that define regex patterns, priority scores, and context words for each national identifier.**

OpenMed detects and sanitizes clinical identifiers—such as Brazilian **CPF/CNPJ**, Dutch **BSN**, and other national ID formats—through a modular internationalization system. The detection logic resides in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py), where each identifier type combines a regular expression pattern with a checksum validator. This architecture allows you to extend or modify ID detection without retraining the underlying NER model.

## Understanding the Two-Step Detection Mechanism

OpenMed identifies clinical IDs using a two-layer approach implemented in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py).

**1. Validator functions** are pure Python functions that verify format and checksum. They follow the signature `def validate_<locale>_<type>(text: str) -> bool`. For example, `validate_portuguese_cpf` strips non-numeric characters, verifies the 11-digit length, and executes the modulo-11 checksum algorithm (see lines [14‑46](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/pii_i18n.py#L14-L46)).

**2. `PIIPattern` objects** wrap regex patterns with metadata. Each pattern includes:
- A compiled regex (e.g., `r"\b\d{3}\.?\d{3}\.?\d{3}-?\d{2}\b"` for CPF)
- The label `"national_id"`
- **Priority** (higher values match earlier)
- **Base score** and **context boost** for confidence scoring
- **Context words** (e.g., `["cpf", "cadastro de pessoas físicas"]`) that increase match likelihood when found nearby
- A reference to the validator function

These patterns are organized into language-specific lists (e.g., `_PORTUGUESE_PII_PATTERNS`, `_DUTCH_PII_PATTERNS`) and collected into the `LANGUAGE_PII_PATTERNS` dictionary at the bottom of the file.

## Step-by-Step Guide to Adding or Modifying ID Providers

### Step 1: Add or Modify a Validator Function

Create a new validator or edit an existing one in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py). The function must accept a string and return a boolean indicating validity.

```python
def validate_xyz_id(text: str) -> bool:
    """Validate a 12-digit XYZ ID using the Luhn algorithm."""
    import re
    digits = re.sub(r"[^\d]", "", text)
    if len(digits) != 12:
        return False
    
    total = 0
    rev = digits[::-1]
    for i, d in enumerate(rev):
        n = int(d)
        if i % 2 == 1:
            n *= 2
            if n > 9:
                n -= 9
        total += n
    return total % 10 == 0

```

Place this function alongside existing validators like `validate_portuguese_cpf` and `validate_dutch_bsn`.

### Step 2: Create a PIIPattern Entry

Define a `PIIPattern` inside the appropriate language block (e.g., `_ENGLISH_PII_PATTERNS` or `_PORTUGUESE_PII_PATTERNS`). Reference your validator in the `validator` parameter.

```python
from openmed.core.pii_i18n import PIIPattern

_ENGLISH_PII_PATTERNS = [
    # existing patterns...

    PIIPattern(
        r"\b\d{3}-\d{3}-\d{3}-\d{3}\b",          # e.g., 123-456-789-012

        "national_id",
        priority=10,
        base_score=0.5,
        context_words=["xyz", "xyz-id", "identifier"],
        context_boost=0.5,
        validator=validate_xyz_id,
    ),
]

```

Omit the `validator` parameter if you only need regex matching without checksum verification.

### Step 3: Register the Pattern in the Language Map

Ensure your language block is referenced in the `LANGUAGE_PII_PATTERNS` dictionary at the bottom of [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py):

```python
LANGUAGE_PII_PATTERNS = {
    "en": _ENGLISH_PII_PATTERNS,
    "pt": _PORTUGUESE_PII_PATTERNS,
    "nl": _DUTCH_PII_PATTERNS,
    # ... other locales

}

```

The NER pipeline consumes this mapping via [`openmed/ner/adapter.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/adapter.py) to apply patterns at inference time.

### Step 4: Validate with Unit Tests

Add tests to [`tests/unit/test_pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_pii_i18n.py) to verify your validator and pattern integration:

```python
def test_validate_xyz_id():
    assert validate_xyz_id("123-456-789-012")
    assert not validate_xyz_id("123-456-789-013")

def test_xyz_id_pattern():
    from openmed.core.pii_i18n import LANGUAGE_PII_PATTERNS
    pattern = next(p for p in LANGUAGE_PII_PATTERNS["en"] 
                   if "xyz" in str(p.context_words))
    assert pattern.validator("123-456-789-012")

```

Run `pytest tests/unit/test_pii_i18n.py` to confirm the integration works correctly.

## Code Examples for Common Customizations

### Adding a Custom Validator for a New National ID

This example implements a complete 12-digit ID validator with Luhn checksum and registers it for English clinical notes:

```python

# openmed/core/pii_i18n.py

def validate_xyz_id(text: str) -> bool:
    """Validate XYZ ID format: 12 digits with Luhn check."""
    import re
    digits = re.sub(r"\D", "", text)
    if len(digits) != 12:
        return False
    
    checksum = 0
    for i, d in enumerate(reversed(digits)):
        n = int(d)
        if i % 2:
            n *= 2
            if n > 9:
                n -= 9
        checksum += n
    return checksum % 10 == 0

# Add to _ENGLISH_PII_PATTERNS list

_ENGLISH_PII_PATTERNS.append(
    PIIPattern(
        r"\b\d{3}-\d{3}-\d{3}-\d{3}\b",
        "national_id",
        priority=10,
        base_score=0.5,
        context_words=["xyz", "xyz-id", "patient id"],
        context_boost=0.4,
        validator=validate_xyz_id,
    )
)

```

### Tightening Existing Regex Patterns

To reduce false positives on CPF detection, modify `_PORTUGUESE_PII_PATTERNS` to require standard formatting:

```python

# openmed/core/pii_i18n.py

_PORTUGUESE_PII_PATTERNS = [
    PIIPattern(
        r"\b\d{3}\.\d{3}\.\d{3}-\d{2}\b",    # enforce dots and hyphen

        "national_id",
        priority=10,
        base_score=0.45,
        context_words=["cpf", "cadastro"],
        context_boost=0.45,
        validator=validate_portuguese_cpf,
    ),
    # ... other patterns

]

```

This stricter regex (`\d{3}\.\d{3}\.\d{3}-\d{2}`) only matches CPF strings with standard punctuation, ignoring plain 11-digit sequences that might occur in other contexts.

## Key Files and Their Roles

| File | Purpose | Key Components |
|------|---------|----------------|
| [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py) | Contains all validators and `PIIPattern` definitions | `validate_portuguese_cpf`, `validate_dutch_bsn`, `LANGUAGE_PII_PATTERNS` |
| [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py) | Maps generic labels to surrogate replacement tokens | `"CPF": "ID_NUM"` mappings |
| [`openmed/ner/adapter.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/adapter.py) | Loads patterns into the NER pipeline | Consumes `LANGUAGE_PII_PATTERNS` |
| [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py) | Optional runtime configuration | Expose pattern toggles via CLI |
| [`tests/unit/test_pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_pii_i18n.py) | Unit tests for validators | Template for new provider tests |

## Summary

- **Validator functions** in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py) handle checksum logic and format verification for clinical IDs like CPF, CNPJ, and BSN.
- **`PIIPattern` objects** link regex patterns to validators and include scoring metadata (`priority`, `base_score`, `context_boost`) to reduce false positives.
- **Language-specific blocks** (e.g., `_PORTUGUESE_PII_PATTERNS`) organize patterns by locale, aggregated into `LANGUAGE_PII_PATTERNS`.
- **No model retraining** is required for regex-based ID detection; changes take effect at inference time.
- **Unit tests** in [`tests/unit/test_pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_pii_i18n.py) ensure new validators handle edge cases and invalid checksums correctly.

## Frequently Asked Questions

### What is the difference between a validator and a PIIPattern?

A **validator** is a Python function that performs mathematical verification (e.g., checksum calculation) on a candidate string, returning `True` or `False`. A **`PIIPattern`** is a data structure that contains the regex to find candidate strings, metadata for scoring, and an optional reference to the validator. The NER pipeline uses the pattern to locate potential IDs, then invokes the validator to confirm legitimacy.

### Do I need to retrain the NER model after adding a new ID pattern?

No. Because OpenMed consumes `LANGUAGE_PII_PATTERNS` at inference time through [`openmed/ner/adapter.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/adapter.py), adding or modifying patterns in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py) takes effect immediately without retraining. This applies to all regex-based clinical ID providers.

### Can I use the same validator for multiple languages?

Yes. Validators are pure functions independent of language context. You can reference the same validator function in multiple language blocks (e.g., `_PORTUGUESE_PII_PATTERNS` and `_ENGLISH_PII_PATTERNS`) if the same ID format appears in clinical notes across different locales. Simply import the function and assign it to the `validator` parameter in each `PIIPattern` definition.

### How do context words improve detection accuracy?

Context words (configured via the `context_words` list and `context_boost` value) increase the confidence score when a regex match appears near specific terms. For example, a CPF pattern match receives a higher confidence score if the surrounding text contains "cpf" or "cadastro de pessoas físicas". This contextual scoring reduces false positives on random 11-digit numbers that happen to match the CPF regex but appear in unrelated contexts.