# Using the Medical Tokenizer for Custom Tokenization in OpenMed

> Learn to use the medical tokenizer for custom tokenization in OpenMed. Split clinical text into stable, human-readable spans by enabling use_medical_tokenizer and providing an exceptions list.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-12

---

**Enable `use_medical_tokenizer` in `OpenMedConfig` and optionally supply an exceptions list to `medical_tokenize` to split clinical text into stable, human-readable spans without altering the model's input tokenization.**

Using the medical tokenizer for custom tokenization in `maziyarpanahi/openmed` lets you transform raw clinical notes into predictable tokens such as `COVID-19`, `CAR-T`, and `39.8C`. This guide walks through the rule-based tokenizer implemented in [`openmed/processing/tokenization.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/tokenization.py), shows how to toggle it in [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py), and explains how to preserve user-specified phrases while remapping model predictions onto clean output spans.

## Using the Medical Tokenizer for Custom Tokenization: Configuration and Pipeline

The medical tokenizer is entirely independent of the model's internal sub-word tokenization. It is activated through the boolean field `use_medical_tokenizer` defined in [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py) (lines 71-75). When this flag is set, the top-level entry point in [`openmed/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/__init__.py) (lines 73-92) automatically executes two operations after inference:

1. **Clinical tokenization** — `medical_tokenize` splits the raw string into `SpanToken` objects, each carrying token text and exact character offsets.
2. **Prediction remapping** — `remap_predictions_to_tokens` aligns the model's character-level spans to these stable tokens and merges contiguous tokens that share the same label.

Because the tokenizer lives on the output side, it only affects entity spans, confidence scores, and metadata. It is safe to enable on any model without retraining or changing its inputs.

## How `medical_tokenize` and Remapping Work Internally

The core logic resides in [`openmed/processing/tokenization.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/tokenization.py). Between lines 45 and 90, `medical_tokenize` applies a compiled regex stored in `_MEDICAL_TOKEN_PATTERN` to generate `SpanToken` instances. It also accepts an optional *exceptions* list: any phrase in this list is forced to remain a single token regardless of regex matches.

Once the model emits predictions, `remap_predictions_to_tokens` (lines 92-168 in the same file) performs the alignment. For each medical token, it selects the highest-scoring label and merges adjacent tokens with identical labels, using a configurable gap threshold.

Default clinical exceptions are stored in the constant `DEFAULT_MEDICAL_EXCEPTIONS` at the top of the file. This built-in list covers terms such as `COVID-19`, `CAR-T`, and numeric temperature patterns.

## Customizing Tokenization with Exception Phrases

You can tailor the splitter in two ways:

- **Pass a custom list directly** to the `exceptions` argument of `medical_tokenize`.
- **Set `medical_tokenizer_exceptions`** in `OpenMedConfig` or export the environment variable `OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS`.

Both methods merge your phrases with `DEFAULT_MEDICAL_EXCEPTIONS`, so standard clinical terms stay intact automatically.

## Direct Tokenization with `medical_tokenize`

For offline preprocessing or bespoke pipelines, call the tokenizer directly:

```python
from openmed.processing.tokenization import medical_tokenize, SpanToken

text = "Patient received CAR-T therapy on 2022-03-15 and temperature rose to 39.8C."
custom_exceptions = ["2022-03-15"]

tokens: list[SpanToken] = medical_tokenize(text, exceptions=custom_exceptions)

for t in tokens:
    print(f"Token: '{t.text}'  [{t.start}:{t.end}]")

```

**Output:**

```text
Token: 'Patient'          [0:7]
Token: 'received'         [8:16]
Token: 'CAR-T'            [17:22]
Token: 'therapy'          [23:30]
Token: 'on'               [31:33]
Token: '2022-03-15'       [34:44]
Token: 'and'              [45:48]
Token: 'temperature'      [49:60]
Token: 'rose'             [61:65]
Token: 'to'               [66:68]
Token: '39.8C'            [69:74]
Token: '.'                [74:75]

```

The date `2022-03-15` is kept whole because it was supplied as a custom exception, while `CAR-T` and `39.8C` survive thanks to the default exception list.

## Enabling Global Remapping in `analyze_text`

To activate the medical tokenizer inside the standard `analyze_text` call, instantiate `OpenMedConfig` with `use_medical_tokenizer=True`:

```python
from openmed import OpenMedConfig, analyze_text

cfg = OpenMedConfig(
    use_medical_tokenizer=True,
    medical_tokenizer_exceptions=["2022-03-15"]
)

result = analyze_text(
    "Patient received CAR-T therapy on 2022-03-15 and temperature rose to 39.8C.",
    config=cfg,
    model_name="some-ner-model"
)

```

The pipeline runs inference with the model's normal tokenization, then remaps predictions onto the stable medical tokens. The returned metadata contains `"medical_tokenizer": true`, signaling that remapping occurred.

## Integrating `TokenizationHelper` in Custom Pipelines

If you handle post-processing yourself, reuse `TokenizationHelper` to align model inputs and later remap predictions:

```python
from openmed.processing.tokenization import (
    TokenizationHelper,
    medical_tokenize,
    remap_predictions_to_tokens
)
from transformers import AutoTokenizer

hf_tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
helper = TokenizationHelper(tokenizer=hf_tokenizer)

# 1. Tokenize with alignment (model side)

enc = helper.tokenize_with_alignment("Patient has fever 39.8C.")

# 2. Example model predictions (character spans)

predictions = [
    {"start": 13, "end": 18, "entity_group": "SYMPTOM", "score": 0.96}
]

# 3. Medical tokenizer tokens

med_tokens = medical_tokenize("Patient has fever 39.8C.", exceptions=None)

# 4. Remap predictions onto medical tokens

remapped = remap_predictions_to_tokens(
    predictions, "Patient has fever 39.8C.", med_tokens
)
print(remapped)

```

This approach decouples model-side tokenization from output-side clinical tokenization, so you can swap custom tokenizers without reloading the model.

## Setting Exceptions via Environment Variables

You can also inject exceptions without changing code. Export a comma-separated list before starting your script:

```bash
export OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS="2022-03-15,2022-04-01"
python my_script.py

```

`OpenMedConfig` parses this variable automatically and merges the values with `DEFAULT_MEDICAL_EXCEPTIONS`. Every subsequent `analyze_text` call will treat those strings as single tokens.

## Summary

- The medical tokenizer is an **output-side, rule-based processor** in [`openmed/processing/tokenization.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/tokenization.py) that does not interfere with model inputs.
- Enable it globally by setting **`use_medical_tokenizer=True`** in `OpenMedConfig` as implemented in [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py).
- **`medical_tokenize`** splits text into `SpanToken` instances using `_MEDICAL_TOKEN_PATTERN` and an optional exceptions list.
- **`remap_predictions_to_tokens`** aligns character-level predictions to those tokens and merges neighboring same-label spans.
- Customize behavior by passing exceptions directly, setting `medical_tokenizer_exceptions` in Python, or exporting **`OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS`**.

## Frequently Asked Questions

### Does the medical tokenizer change how the model sees input text?

No. According to the `maziyarpanahi/openmed` source code, the tokenizer runs strictly after inference inside [`openmed/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/__init__.py). It remaps output spans, confidence scores, and metadata only, leaving the model's sub-word tokenization completely unchanged.

### How do I force a multi-word phrase to remain a single token?

Pass the phrase to the `exceptions` parameter of `medical_tokenize`, or add it to `OpenMedConfig.medical_tokenizer_exceptions`. Alternatively, set the `OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS` environment variable so the phrase is preserved across all pipeline runs.

### What is the difference between `medical_tokenize` and `remap_predictions_to_tokens`?

`medical_tokenize` performs the clinical text splitting and returns a list of `SpanToken` objects. `remap_predictions_to_tokens` takes the model's raw character-span predictions and aligns them to those tokens, choosing the highest-scoring label per token and merging adjacent tokens that share the same label.

### Is there a built-in list of clinical exceptions?

Yes. The constant `DEFAULT_MEDICAL_EXCEPTIONS` defined at the top of [`openmed/processing/tokenization.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/tokenization.py) includes common clinical strings such as `COVID-19`, `CAR-T`, and numeric temperature values. Any custom exceptions you provide are merged with this default list.