How to Implement Custom Label Normalization with CANONICAL_LABELS in OpenMed
You can customize label normalization in OpenMed by extending the _ALIAS_MAP dictionary or wrapping the normalize_label function in openmed/core/labels.py, ensuring all entity labels map to the canonical UPPER_SNAKE_CASE identifiers defined in the frozen set CANONICAL_LABELS.
OpenMed unifies privacy-preserving entity recognition under a single taxonomy stored in CANONICAL_LABELS. Regardless of whether models output English labels, Portuguese tags, or BIOES-prefixed tokens, the normalize_label function acts as the central gateway that resolves every variation to a standardized identifier. This design allows you to implement custom label normalization without touching downstream components like the anonymizer or model registry.
Understanding the Label Normalization Pipeline
The normalization logic lives entirely within openmed/core/labels.py. The public entry point normalize_label(label: str, lang: str = "en") → str (lines 71-92) processes incoming strings through a strict four-step workflow before returning a canonical identifier.
The CANONICAL_LABELS Taxonomy
According to the source code at lines 99-112, OpenMed defines its complete label vocabulary as a frozen set of UPPER_SNAKE_CASE strings. This set includes standardized identifiers such as FIRST_NAME, ID_NUM, DATE_OF_BIRTH, and PHONE. By freezing this collection, OpenMed ensures that all downstream modules—from the anonymizer to the model registry—operate on a stable, immutable taxonomy.
The Four-Step Normalization Process
The normalize_label function executes the following pipeline as implemented in openmed/core/labels.py:
- Sanitization: The internal
_keyhelper (lines 65-69) converts input to lowercase, strips non-alphanumeric characters, and removes BIOES prefixes (B-,I-,E-,S-). - Alias lookup: The cleaned token is matched against
_ALIAS_MAP(lines 19-53), a dictionary mapping normalized tokens to canonical labels. - Fallback matching: If no alias exists, the code attempts direct upper-case matching against
CANONICAL_LABELSafter converting hyphens and spaces to underscores (lines 104-108). - Default resolution: Any unknown label resolves to
OTHER.
Extending the Alias Map for Custom Labels
The simplest way to support new label styles is to extend the _ALIAS_MAP dictionary. This approach requires no changes to the normalize_label function itself because the function checks this map before attempting fallback matching.
For example, to support a model that emits "patient_id" or "mrn":
# openmed/core/labels.py
_ALIAS_MAP: Final[Mapping[str, str]] = {
# ... existing entries ...
"patientid": ID_NUM, # maps "patient_id" -> ID_NUM
"mrn": ID_NUM, # medical record number alias
}
After this modification, the function correctly normalizes variations:
from openmed.core.labels import normalize_label
print(normalize_label("patient_id")) # Output: ID_NUM
print(normalize_label("B-MRN")) # Output: ID_NUM (BIOES prefix stripped)
Creating Custom Normalization Wrappers
For complex logic such as language-specific disambiguation, wrap the core function rather than modifying the alias map. This pattern preserves the original implementation while injecting custom rules that execute before the standard lookup.
The following example handles Portuguese-specific labels that require special treatment:
# my_project/custom_normalisation.py
from openmed.core.labels import normalize_label, OTHER
def normalize_label_pt(label: str) -> str:
"""Portuguese-specific tweaks before delegating to the library."""
# Portuguese models sometimes emit "nome" for generic names
if label.lower() == "nome":
return "PERSON"
return normalize_label(label, lang="pt")
Usage demonstrates the fallback behavior:
from my_project.custom_normalisation import normalize_label_pt
print(normalize_label_pt("nome")) # Output: PERSON (custom rule)
print(normalize_label_pt("telefone")) # Output: PHONE (falls through to alias map)
Integrating Custom Logic into the Anonymizer
Downstream components like the anonymizer engine consume normalize_label to select appropriate fake data generators. You can inject your custom wrapper by importing it in place of the standard function.
In openmed/core/anonymizer/engine.py, modify the import to use your custom implementation:
# openmed/core/anonymizer/engine.py (excerpt)
from my_project.custom_normalisation import normalize_label_pt as normalize_label
def _anonymise(self, label: str, value: str) -> str:
canonical = normalize_label(label) # Uses custom Portuguese logic
generator = LABEL_GENERATORS.get(canonical, LABEL_GENERATORS["OTHER"])
# ... generate fake value ...
This integration ensures that all labels pass through your custom normalization before the anonymization pipeline processes them, while the rest of the OpenMed codebase remains unchanged.
Summary
- OpenMed maintains a single taxonomy in the
CANONICAL_LABELSfrozen set defined inopenmed/core/labels.py(lines 99-112). - The
normalize_labelfunction provides a four-step pipeline: sanitization via_key, alias lookup in_ALIAS_MAP, fallback matching againstCANONICAL_LABELS, and default resolution toOTHER. - Extend
_ALIAS_MAP(lines 19-53) to add simple string mappings for new label variations without modifying function logic. - Create wrapper functions around
normalize_labelfor complex logic like language-specific disambiguation. - The anonymizer engine and other downstream components automatically consume the canonical output, requiring no modifications when you customize the normalizer.
Frequently Asked Questions
What is the purpose of CANONICAL_LABELS in OpenMed?
CANONICAL_LABELS is a frozen set of UPPER_SNAKE_CASE identifiers that serves as the single taxonomy for all privacy-preserving entities in OpenMed. It ensures that models emitting different label formats—whether English, Portuguese, or BIOES-tagged—all resolve to standardized names like FIRST_NAME or ID_NUM, enabling consistent downstream processing across the entire pipeline.
How does the _ALIAS_MAP dictionary work?
_ALIAS_MAP is a dictionary defined at lines 19-53 of openmed/core/labels.py that maps normalized tokens (lowercase, stripped of punctuation) to canonical labels. When normalize_label receives input, it cleans the string using the _key helper and looks up the result in this map. If found, it returns the corresponding canonical label immediately without checking other fallbacks.
Can I override normalize_label without modifying the source code?
Yes. You can create a wrapper function in your own module that imports and calls the original normalize_label from openmed/core/labels.py, applying your custom logic before or after the call. Import your wrapper instead of the original function in your application entry points or configuration files. This approach avoids forking the repository while still customizing the normalization behavior for your specific use case.
Where are the label normalization tests located?
Unit tests verifying the alias mapping and canonical label round-tripping reside in tests/unit/core/test_labels.py. Additional consistency checks ensuring model-specific label sets map correctly to the canonical taxonomy are located in tests/unit/ner/test_label_map_consistency.py. These files validate that all custom aliases resolve to members of the CANONICAL_LABELS set.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →