How to Customize Clinical ID Providers (CPF, CNPJ, BSN, etc.) in OpenMed
Customize clinical ID providers in OpenMed by implementing validator functions in openmed/core/pii_i18n.py and linking them to PIIPattern objects that define regex patterns, priority scores, and context words for each national identifier.
OpenMed detects and sanitizes clinical identifiers—such as Brazilian CPF/CNPJ, Dutch BSN, and other national ID formats—through a modular internationalization system. The detection logic resides in openmed/core/pii_i18n.py, where each identifier type combines a regular expression pattern with a checksum validator. This architecture allows you to extend or modify ID detection without retraining the underlying NER model.
Understanding the Two-Step Detection Mechanism
OpenMed identifies clinical IDs using a two-layer approach implemented in openmed/core/pii_i18n.py.
1. Validator functions are pure Python functions that verify format and checksum. They follow the signature def validate_<locale>_<type>(text: str) -> bool. For example, validate_portuguese_cpf strips non-numeric characters, verifies the 11-digit length, and executes the modulo-11 checksum algorithm (see lines 14‑46).
2. PIIPattern objects wrap regex patterns with metadata. Each pattern includes:
- A compiled regex (e.g.,
r"\b\d{3}\.?\d{3}\.?\d{3}-?\d{2}\b"for CPF) - The label
"national_id" - Priority (higher values match earlier)
- Base score and context boost for confidence scoring
- Context words (e.g.,
["cpf", "cadastro de pessoas físicas"]) that increase match likelihood when found nearby - A reference to the validator function
These patterns are organized into language-specific lists (e.g., _PORTUGUESE_PII_PATTERNS, _DUTCH_PII_PATTERNS) and collected into the LANGUAGE_PII_PATTERNS dictionary at the bottom of the file.
Step-by-Step Guide to Adding or Modifying ID Providers
Step 1: Add or Modify a Validator Function
Create a new validator or edit an existing one in openmed/core/pii_i18n.py. The function must accept a string and return a boolean indicating validity.
def validate_xyz_id(text: str) -> bool:
"""Validate a 12-digit XYZ ID using the Luhn algorithm."""
import re
digits = re.sub(r"[^\d]", "", text)
if len(digits) != 12:
return False
total = 0
rev = digits[::-1]
for i, d in enumerate(rev):
n = int(d)
if i % 2 == 1:
n *= 2
if n > 9:
n -= 9
total += n
return total % 10 == 0
Place this function alongside existing validators like validate_portuguese_cpf and validate_dutch_bsn.
Step 2: Create a PIIPattern Entry
Define a PIIPattern inside the appropriate language block (e.g., _ENGLISH_PII_PATTERNS or _PORTUGUESE_PII_PATTERNS). Reference your validator in the validator parameter.
from openmed.core.pii_i18n import PIIPattern
_ENGLISH_PII_PATTERNS = [
# existing patterns...
PIIPattern(
r"\b\d{3}-\d{3}-\d{3}-\d{3}\b", # e.g., 123-456-789-012
"national_id",
priority=10,
base_score=0.5,
context_words=["xyz", "xyz-id", "identifier"],
context_boost=0.5,
validator=validate_xyz_id,
),
]
Omit the validator parameter if you only need regex matching without checksum verification.
Step 3: Register the Pattern in the Language Map
Ensure your language block is referenced in the LANGUAGE_PII_PATTERNS dictionary at the bottom of openmed/core/pii_i18n.py:
LANGUAGE_PII_PATTERNS = {
"en": _ENGLISH_PII_PATTERNS,
"pt": _PORTUGUESE_PII_PATTERNS,
"nl": _DUTCH_PII_PATTERNS,
# ... other locales
}
The NER pipeline consumes this mapping via openmed/ner/adapter.py to apply patterns at inference time.
Step 4: Validate with Unit Tests
Add tests to tests/unit/test_pii_i18n.py to verify your validator and pattern integration:
def test_validate_xyz_id():
assert validate_xyz_id("123-456-789-012")
assert not validate_xyz_id("123-456-789-013")
def test_xyz_id_pattern():
from openmed.core.pii_i18n import LANGUAGE_PII_PATTERNS
pattern = next(p for p in LANGUAGE_PII_PATTERNS["en"]
if "xyz" in str(p.context_words))
assert pattern.validator("123-456-789-012")
Run pytest tests/unit/test_pii_i18n.py to confirm the integration works correctly.
Code Examples for Common Customizations
Adding a Custom Validator for a New National ID
This example implements a complete 12-digit ID validator with Luhn checksum and registers it for English clinical notes:
# openmed/core/pii_i18n.py
def validate_xyz_id(text: str) -> bool:
"""Validate XYZ ID format: 12 digits with Luhn check."""
import re
digits = re.sub(r"\D", "", text)
if len(digits) != 12:
return False
checksum = 0
for i, d in enumerate(reversed(digits)):
n = int(d)
if i % 2:
n *= 2
if n > 9:
n -= 9
checksum += n
return checksum % 10 == 0
# Add to _ENGLISH_PII_PATTERNS list
_ENGLISH_PII_PATTERNS.append(
PIIPattern(
r"\b\d{3}-\d{3}-\d{3}-\d{3}\b",
"national_id",
priority=10,
base_score=0.5,
context_words=["xyz", "xyz-id", "patient id"],
context_boost=0.4,
validator=validate_xyz_id,
)
)
Tightening Existing Regex Patterns
To reduce false positives on CPF detection, modify _PORTUGUESE_PII_PATTERNS to require standard formatting:
# openmed/core/pii_i18n.py
_PORTUGUESE_PII_PATTERNS = [
PIIPattern(
r"\b\d{3}\.\d{3}\.\d{3}-\d{2}\b", # enforce dots and hyphen
"national_id",
priority=10,
base_score=0.45,
context_words=["cpf", "cadastro"],
context_boost=0.45,
validator=validate_portuguese_cpf,
),
# ... other patterns
]
This stricter regex (\d{3}\.\d{3}\.\d{3}-\d{2}) only matches CPF strings with standard punctuation, ignoring plain 11-digit sequences that might occur in other contexts.
Key Files and Their Roles
| File | Purpose | Key Components |
|---|---|---|
openmed/core/pii_i18n.py |
Contains all validators and PIIPattern definitions |
validate_portuguese_cpf, validate_dutch_bsn, LANGUAGE_PII_PATTERNS |
openmed/core/pii.py |
Maps generic labels to surrogate replacement tokens | "CPF": "ID_NUM" mappings |
openmed/ner/adapter.py |
Loads patterns into the NER pipeline | Consumes LANGUAGE_PII_PATTERNS |
openmed/core/config.py |
Optional runtime configuration | Expose pattern toggles via CLI |
tests/unit/test_pii_i18n.py |
Unit tests for validators | Template for new provider tests |
Summary
- Validator functions in
openmed/core/pii_i18n.pyhandle checksum logic and format verification for clinical IDs like CPF, CNPJ, and BSN. PIIPatternobjects link regex patterns to validators and include scoring metadata (priority,base_score,context_boost) to reduce false positives.- Language-specific blocks (e.g.,
_PORTUGUESE_PII_PATTERNS) organize patterns by locale, aggregated intoLANGUAGE_PII_PATTERNS. - No model retraining is required for regex-based ID detection; changes take effect at inference time.
- Unit tests in
tests/unit/test_pii_i18n.pyensure new validators handle edge cases and invalid checksums correctly.
Frequently Asked Questions
What is the difference between a validator and a PIIPattern?
A validator is a Python function that performs mathematical verification (e.g., checksum calculation) on a candidate string, returning True or False. A PIIPattern is a data structure that contains the regex to find candidate strings, metadata for scoring, and an optional reference to the validator. The NER pipeline uses the pattern to locate potential IDs, then invokes the validator to confirm legitimacy.
Do I need to retrain the NER model after adding a new ID pattern?
No. Because OpenMed consumes LANGUAGE_PII_PATTERNS at inference time through openmed/ner/adapter.py, adding or modifying patterns in openmed/core/pii_i18n.py takes effect immediately without retraining. This applies to all regex-based clinical ID providers.
Can I use the same validator for multiple languages?
Yes. Validators are pure functions independent of language context. You can reference the same validator function in multiple language blocks (e.g., _PORTUGUESE_PII_PATTERNS and _ENGLISH_PII_PATTERNS) if the same ID format appears in clinical notes across different locales. Simply import the function and assign it to the validator parameter in each PIIPattern definition.
How do context words improve detection accuracy?
Context words (configured via the context_words list and context_boost value) increase the confidence score when a regex match appears near specific terms. For example, a CPF pattern match receives a higher confidence score if the surrounding text contains "cpf" or "cadastro de pessoas físicas". This contextual scoring reduces false positives on random 11-digit numbers that happen to match the CPF regex but appear in unrelated contexts.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →