Using the Medical Tokenizer for Custom Tokenization in OpenMed

Enable use_medical_tokenizer in OpenMedConfig and optionally supply an exceptions list to medical_tokenize to split clinical text into stable, human-readable spans without altering the model's input tokenization.

Using the medical tokenizer for custom tokenization in maziyarpanahi/openmed lets you transform raw clinical notes into predictable tokens such as COVID-19, CAR-T, and 39.8C. This guide walks through the rule-based tokenizer implemented in openmed/processing/tokenization.py, shows how to toggle it in openmed/core/config.py, and explains how to preserve user-specified phrases while remapping model predictions onto clean output spans.

Using the Medical Tokenizer for Custom Tokenization: Configuration and Pipeline

The medical tokenizer is entirely independent of the model's internal sub-word tokenization. It is activated through the boolean field use_medical_tokenizer defined in openmed/core/config.py (lines 71-75). When this flag is set, the top-level entry point in openmed/__init__.py (lines 73-92) automatically executes two operations after inference:

  1. Clinical tokenization — medical_tokenize splits the raw string into SpanToken objects, each carrying token text and exact character offsets.
  2. Prediction remapping — remap_predictions_to_tokens aligns the model's character-level spans to these stable tokens and merges contiguous tokens that share the same label.

Because the tokenizer lives on the output side, it only affects entity spans, confidence scores, and metadata. It is safe to enable on any model without retraining or changing its inputs.

How medical_tokenize and Remapping Work Internally

The core logic resides in openmed/processing/tokenization.py. Between lines 45 and 90, medical_tokenize applies a compiled regex stored in _MEDICAL_TOKEN_PATTERN to generate SpanToken instances. It also accepts an optional exceptions list: any phrase in this list is forced to remain a single token regardless of regex matches.

Once the model emits predictions, remap_predictions_to_tokens (lines 92-168 in the same file) performs the alignment. For each medical token, it selects the highest-scoring label and merges adjacent tokens with identical labels, using a configurable gap threshold.

Default clinical exceptions are stored in the constant DEFAULT_MEDICAL_EXCEPTIONS at the top of the file. This built-in list covers terms such as COVID-19, CAR-T, and numeric temperature patterns.

Customizing Tokenization with Exception Phrases

You can tailor the splitter in two ways:

  • Pass a custom list directly to the exceptions argument of medical_tokenize.
  • Set medical_tokenizer_exceptions in OpenMedConfig or export the environment variable OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS.

Both methods merge your phrases with DEFAULT_MEDICAL_EXCEPTIONS, so standard clinical terms stay intact automatically.

Direct Tokenization with medical_tokenize

For offline preprocessing or bespoke pipelines, call the tokenizer directly:

from openmed.processing.tokenization import medical_tokenize, SpanToken

text = "Patient received CAR-T therapy on 2022-03-15 and temperature rose to 39.8C."
custom_exceptions = ["2022-03-15"]

tokens: list[SpanToken] = medical_tokenize(text, exceptions=custom_exceptions)

for t in tokens:
    print(f"Token: '{t.text}'  [{t.start}:{t.end}]")

Output:

Token: 'Patient'          [0:7]
Token: 'received'         [8:16]
Token: 'CAR-T'            [17:22]
Token: 'therapy'          [23:30]
Token: 'on'               [31:33]
Token: '2022-03-15'       [34:44]
Token: 'and'              [45:48]
Token: 'temperature'      [49:60]
Token: 'rose'             [61:65]
Token: 'to'               [66:68]
Token: '39.8C'            [69:74]
Token: '.'                [74:75]

The date 2022-03-15 is kept whole because it was supplied as a custom exception, while CAR-T and 39.8C survive thanks to the default exception list.

Enabling Global Remapping in analyze_text

To activate the medical tokenizer inside the standard analyze_text call, instantiate OpenMedConfig with use_medical_tokenizer=True:

from openmed import OpenMedConfig, analyze_text

cfg = OpenMedConfig(
    use_medical_tokenizer=True,
    medical_tokenizer_exceptions=["2022-03-15"]
)

result = analyze_text(
    "Patient received CAR-T therapy on 2022-03-15 and temperature rose to 39.8C.",
    config=cfg,
    model_name="some-ner-model"
)

The pipeline runs inference with the model's normal tokenization, then remaps predictions onto the stable medical tokens. The returned metadata contains "medical_tokenizer": true, signaling that remapping occurred.

Integrating TokenizationHelper in Custom Pipelines

If you handle post-processing yourself, reuse TokenizationHelper to align model inputs and later remap predictions:

from openmed.processing.tokenization import (
    TokenizationHelper,
    medical_tokenize,
    remap_predictions_to_tokens
)
from transformers import AutoTokenizer

hf_tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
helper = TokenizationHelper(tokenizer=hf_tokenizer)

# 1. Tokenize with alignment (model side)

enc = helper.tokenize_with_alignment("Patient has fever 39.8C.")

# 2. Example model predictions (character spans)

predictions = [
    {"start": 13, "end": 18, "entity_group": "SYMPTOM", "score": 0.96}
]

# 3. Medical tokenizer tokens

med_tokens = medical_tokenize("Patient has fever 39.8C.", exceptions=None)

# 4. Remap predictions onto medical tokens

remapped = remap_predictions_to_tokens(
    predictions, "Patient has fever 39.8C.", med_tokens
)
print(remapped)

This approach decouples model-side tokenization from output-side clinical tokenization, so you can swap custom tokenizers without reloading the model.

Setting Exceptions via Environment Variables

You can also inject exceptions without changing code. Export a comma-separated list before starting your script:

export OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS="2022-03-15,2022-04-01"
python my_script.py

OpenMedConfig parses this variable automatically and merges the values with DEFAULT_MEDICAL_EXCEPTIONS. Every subsequent analyze_text call will treat those strings as single tokens.

Summary

  • The medical tokenizer is an output-side, rule-based processor in openmed/processing/tokenization.py that does not interfere with model inputs.
  • Enable it globally by setting use_medical_tokenizer=True in OpenMedConfig as implemented in openmed/core/config.py.
  • medical_tokenize splits text into SpanToken instances using _MEDICAL_TOKEN_PATTERN and an optional exceptions list.
  • remap_predictions_to_tokens aligns character-level predictions to those tokens and merges neighboring same-label spans.
  • Customize behavior by passing exceptions directly, setting medical_tokenizer_exceptions in Python, or exporting OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS.

Frequently Asked Questions

Does the medical tokenizer change how the model sees input text?

No. According to the maziyarpanahi/openmed source code, the tokenizer runs strictly after inference inside openmed/__init__.py. It remaps output spans, confidence scores, and metadata only, leaving the model's sub-word tokenization completely unchanged.

How do I force a multi-word phrase to remain a single token?

Pass the phrase to the exceptions parameter of medical_tokenize, or add it to OpenMedConfig.medical_tokenizer_exceptions. Alternatively, set the OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS environment variable so the phrase is preserved across all pipeline runs.

What is the difference between medical_tokenize and remap_predictions_to_tokens?

medical_tokenize performs the clinical text splitting and returns a list of SpanToken objects. remap_predictions_to_tokens takes the model's raw character-span predictions and aligns them to those tokens, choosing the highest-scoring label per token and merging adjacent tokens that share the same label.

Is there a built-in list of clinical exceptions?

Yes. The constant DEFAULT_MEDICAL_EXCEPTIONS defined at the top of openmed/processing/tokenization.py includes common clinical strings such as COVID-19, CAR-T, and numeric temperature values. Any custom exceptions you provide are merged with this default list.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →