# How to Implement Custom Label Normalization with CANONICAL_LABELS in OpenMed

> Learn to implement custom label normalization with CANONICAL_LABELS in OpenMed. Extend _ALIAS_MAP or wrap normalize_label for consistent entity labeling.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**You can customize label normalization in OpenMed by extending the `_ALIAS_MAP` dictionary or wrapping the `normalize_label` function in [`openmed/core/labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/labels.py), ensuring all entity labels map to the canonical `UPPER_SNAKE_CASE` identifiers defined in the frozen set `CANONICAL_LABELS`.**

OpenMed unifies privacy-preserving entity recognition under a single taxonomy stored in `CANONICAL_LABELS`. Regardless of whether models output English labels, Portuguese tags, or BIOES-prefixed tokens, the `normalize_label` function acts as the central gateway that resolves every variation to a standardized identifier. This design allows you to implement custom label normalization without touching downstream components like the anonymizer or model registry.

## Understanding the Label Normalization Pipeline

The normalization logic lives entirely within [`openmed/core/labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/labels.py). The public entry point `normalize_label(label: str, lang: str = "en") → str` (lines 71-92) processes incoming strings through a strict four-step workflow before returning a canonical identifier.

### The CANONICAL_LABELS Taxonomy

According to the source code at lines 99-112, OpenMed defines its complete label vocabulary as a frozen set of `UPPER_SNAKE_CASE` strings. This set includes standardized identifiers such as `FIRST_NAME`, `ID_NUM`, `DATE_OF_BIRTH`, and `PHONE`. By freezing this collection, OpenMed ensures that all downstream modules—from the anonymizer to the model registry—operate on a stable, immutable taxonomy.

### The Four-Step Normalization Process

The `normalize_label` function executes the following pipeline as implemented in [`openmed/core/labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/labels.py):

1. **Sanitization**: The internal `_key` helper (lines 65-69) converts input to lowercase, strips non-alphanumeric characters, and removes BIOES prefixes (`B-`, `I-`, `E-`, `S-`).
2. **Alias lookup**: The cleaned token is matched against `_ALIAS_MAP` (lines 19-53), a dictionary mapping normalized tokens to canonical labels.
3. **Fallback matching**: If no alias exists, the code attempts direct upper-case matching against `CANONICAL_LABELS` after converting hyphens and spaces to underscores (lines 104-108).
4. **Default resolution**: Any unknown label resolves to `OTHER`.

## Extending the Alias Map for Custom Labels

The simplest way to support new label styles is to extend the `_ALIAS_MAP` dictionary. This approach requires no changes to the `normalize_label` function itself because the function checks this map before attempting fallback matching.

For example, to support a model that emits `"patient_id"` or `"mrn"`:

```python

# openmed/core/labels.py

_ALIAS_MAP: Final[Mapping[str, str]] = {
    # ... existing entries ...

    "patientid": ID_NUM,          # maps "patient_id" -> ID_NUM

    "mrn": ID_NUM,                # medical record number alias

}

```

After this modification, the function correctly normalizes variations:

```python
from openmed.core.labels import normalize_label

print(normalize_label("patient_id"))  # Output: ID_NUM

print(normalize_label("B-MRN"))        # Output: ID_NUM (BIOES prefix stripped)

```

## Creating Custom Normalization Wrappers

For complex logic such as language-specific disambiguation, wrap the core function rather than modifying the alias map. This pattern preserves the original implementation while injecting custom rules that execute before the standard lookup.

The following example handles Portuguese-specific labels that require special treatment:

```python

# my_project/custom_normalisation.py

from openmed.core.labels import normalize_label, OTHER

def normalize_label_pt(label: str) -> str:
    """Portuguese-specific tweaks before delegating to the library."""
    # Portuguese models sometimes emit "nome" for generic names

    if label.lower() == "nome":
        return "PERSON"
    return normalize_label(label, lang="pt")

```

Usage demonstrates the fallback behavior:

```python
from my_project.custom_normalisation import normalize_label_pt

print(normalize_label_pt("nome"))      # Output: PERSON (custom rule)

print(normalize_label_pt("telefone")) # Output: PHONE (falls through to alias map)

```

## Integrating Custom Logic into the Anonymizer

Downstream components like the anonymizer engine consume `normalize_label` to select appropriate fake data generators. You can inject your custom wrapper by importing it in place of the standard function.

In [`openmed/core/anonymizer/engine.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/anonymizer/engine.py), modify the import to use your custom implementation:

```python

# openmed/core/anonymizer/engine.py (excerpt)

from my_project.custom_normalisation import normalize_label_pt as normalize_label

def _anonymise(self, label: str, value: str) -> str:
    canonical = normalize_label(label)  # Uses custom Portuguese logic

    generator = LABEL_GENERATORS.get(canonical, LABEL_GENERATORS["OTHER"])
    # ... generate fake value ...

```

This integration ensures that all labels pass through your custom normalization before the anonymization pipeline processes them, while the rest of the OpenMed codebase remains unchanged.

## Summary

- OpenMed maintains a single taxonomy in the `CANONICAL_LABELS` frozen set defined in [`openmed/core/labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/labels.py) (lines 99-112).
- The `normalize_label` function provides a four-step pipeline: sanitization via `_key`, alias lookup in `_ALIAS_MAP`, fallback matching against `CANONICAL_LABELS`, and default resolution to `OTHER`.
- Extend `_ALIAS_MAP` (lines 19-53) to add simple string mappings for new label variations without modifying function logic.
- Create wrapper functions around `normalize_label` for complex logic like language-specific disambiguation.
- The anonymizer engine and other downstream components automatically consume the canonical output, requiring no modifications when you customize the normalizer.

## Frequently Asked Questions

### What is the purpose of CANONICAL_LABELS in OpenMed?

`CANONICAL_LABELS` is a frozen set of `UPPER_SNAKE_CASE` identifiers that serves as the single taxonomy for all privacy-preserving entities in OpenMed. It ensures that models emitting different label formats—whether English, Portuguese, or BIOES-tagged—all resolve to standardized names like `FIRST_NAME` or `ID_NUM`, enabling consistent downstream processing across the entire pipeline.

### How does the _ALIAS_MAP dictionary work?

`_ALIAS_MAP` is a dictionary defined at lines 19-53 of [`openmed/core/labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/labels.py) that maps normalized tokens (lowercase, stripped of punctuation) to canonical labels. When `normalize_label` receives input, it cleans the string using the `_key` helper and looks up the result in this map. If found, it returns the corresponding canonical label immediately without checking other fallbacks.

### Can I override normalize_label without modifying the source code?

Yes. You can create a wrapper function in your own module that imports and calls the original `normalize_label` from [`openmed/core/labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/labels.py), applying your custom logic before or after the call. Import your wrapper instead of the original function in your application entry points or configuration files. This approach avoids forking the repository while still customizing the normalization behavior for your specific use case.

### Where are the label normalization tests located?

Unit tests verifying the alias mapping and canonical label round-tripping reside in [`tests/unit/core/test_labels.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/core/test_labels.py). Additional consistency checks ensuring model-specific label sets map correctly to the canonical taxonomy are located in [`tests/unit/ner/test_label_map_consistency.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/ner/test_label_map_consistency.py). These files validate that all custom aliases resolve to members of the `CANONICAL_LABELS` set.