Using the Medical Tokenizer for Custom Tokenization in OpenMed
Enable use_medical_tokenizer in OpenMedConfig and optionally supply an exceptions list to medical_tokenize to split clinical text into stable, human-readable spans without altering the model's input tokenization.
Using the medical tokenizer for custom tokenization in maziyarpanahi/openmed lets you transform raw clinical notes into predictable tokens such as COVID-19, CAR-T, and 39.8C. This guide walks through the rule-based tokenizer implemented in openmed/processing/tokenization.py, shows how to toggle it in openmed/core/config.py, and explains how to preserve user-specified phrases while remapping model predictions onto clean output spans.
Using the Medical Tokenizer for Custom Tokenization: Configuration and Pipeline
The medical tokenizer is entirely independent of the model's internal sub-word tokenization. It is activated through the boolean field use_medical_tokenizer defined in openmed/core/config.py (lines 71-75). When this flag is set, the top-level entry point in openmed/__init__.py (lines 73-92) automatically executes two operations after inference:
- Clinical tokenization —
medical_tokenizesplits the raw string intoSpanTokenobjects, each carrying token text and exact character offsets. - Prediction remapping —
remap_predictions_to_tokensaligns the model's character-level spans to these stable tokens and merges contiguous tokens that share the same label.
Because the tokenizer lives on the output side, it only affects entity spans, confidence scores, and metadata. It is safe to enable on any model without retraining or changing its inputs.
How medical_tokenize and Remapping Work Internally
The core logic resides in openmed/processing/tokenization.py. Between lines 45 and 90, medical_tokenize applies a compiled regex stored in _MEDICAL_TOKEN_PATTERN to generate SpanToken instances. It also accepts an optional exceptions list: any phrase in this list is forced to remain a single token regardless of regex matches.
Once the model emits predictions, remap_predictions_to_tokens (lines 92-168 in the same file) performs the alignment. For each medical token, it selects the highest-scoring label and merges adjacent tokens with identical labels, using a configurable gap threshold.
Default clinical exceptions are stored in the constant DEFAULT_MEDICAL_EXCEPTIONS at the top of the file. This built-in list covers terms such as COVID-19, CAR-T, and numeric temperature patterns.
Customizing Tokenization with Exception Phrases
You can tailor the splitter in two ways:
- Pass a custom list directly to the
exceptionsargument ofmedical_tokenize. - Set
medical_tokenizer_exceptionsinOpenMedConfigor export the environment variableOPENMED_MEDICAL_TOKENIZER_EXCEPTIONS.
Both methods merge your phrases with DEFAULT_MEDICAL_EXCEPTIONS, so standard clinical terms stay intact automatically.
Direct Tokenization with medical_tokenize
For offline preprocessing or bespoke pipelines, call the tokenizer directly:
from openmed.processing.tokenization import medical_tokenize, SpanToken
text = "Patient received CAR-T therapy on 2022-03-15 and temperature rose to 39.8C."
custom_exceptions = ["2022-03-15"]
tokens: list[SpanToken] = medical_tokenize(text, exceptions=custom_exceptions)
for t in tokens:
print(f"Token: '{t.text}' [{t.start}:{t.end}]")
Output:
Token: 'Patient' [0:7]
Token: 'received' [8:16]
Token: 'CAR-T' [17:22]
Token: 'therapy' [23:30]
Token: 'on' [31:33]
Token: '2022-03-15' [34:44]
Token: 'and' [45:48]
Token: 'temperature' [49:60]
Token: 'rose' [61:65]
Token: 'to' [66:68]
Token: '39.8C' [69:74]
Token: '.' [74:75]
The date 2022-03-15 is kept whole because it was supplied as a custom exception, while CAR-T and 39.8C survive thanks to the default exception list.
Enabling Global Remapping in analyze_text
To activate the medical tokenizer inside the standard analyze_text call, instantiate OpenMedConfig with use_medical_tokenizer=True:
from openmed import OpenMedConfig, analyze_text
cfg = OpenMedConfig(
use_medical_tokenizer=True,
medical_tokenizer_exceptions=["2022-03-15"]
)
result = analyze_text(
"Patient received CAR-T therapy on 2022-03-15 and temperature rose to 39.8C.",
config=cfg,
model_name="some-ner-model"
)
The pipeline runs inference with the model's normal tokenization, then remaps predictions onto the stable medical tokens. The returned metadata contains "medical_tokenizer": true, signaling that remapping occurred.
Integrating TokenizationHelper in Custom Pipelines
If you handle post-processing yourself, reuse TokenizationHelper to align model inputs and later remap predictions:
from openmed.processing.tokenization import (
TokenizationHelper,
medical_tokenize,
remap_predictions_to_tokens
)
from transformers import AutoTokenizer
hf_tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
helper = TokenizationHelper(tokenizer=hf_tokenizer)
# 1. Tokenize with alignment (model side)
enc = helper.tokenize_with_alignment("Patient has fever 39.8C.")
# 2. Example model predictions (character spans)
predictions = [
{"start": 13, "end": 18, "entity_group": "SYMPTOM", "score": 0.96}
]
# 3. Medical tokenizer tokens
med_tokens = medical_tokenize("Patient has fever 39.8C.", exceptions=None)
# 4. Remap predictions onto medical tokens
remapped = remap_predictions_to_tokens(
predictions, "Patient has fever 39.8C.", med_tokens
)
print(remapped)
This approach decouples model-side tokenization from output-side clinical tokenization, so you can swap custom tokenizers without reloading the model.
Setting Exceptions via Environment Variables
You can also inject exceptions without changing code. Export a comma-separated list before starting your script:
export OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS="2022-03-15,2022-04-01"
python my_script.py
OpenMedConfig parses this variable automatically and merges the values with DEFAULT_MEDICAL_EXCEPTIONS. Every subsequent analyze_text call will treat those strings as single tokens.
Summary
- The medical tokenizer is an output-side, rule-based processor in
openmed/processing/tokenization.pythat does not interfere with model inputs. - Enable it globally by setting
use_medical_tokenizer=TrueinOpenMedConfigas implemented inopenmed/core/config.py. medical_tokenizesplits text intoSpanTokeninstances using_MEDICAL_TOKEN_PATTERNand an optional exceptions list.remap_predictions_to_tokensaligns character-level predictions to those tokens and merges neighboring same-label spans.- Customize behavior by passing exceptions directly, setting
medical_tokenizer_exceptionsin Python, or exportingOPENMED_MEDICAL_TOKENIZER_EXCEPTIONS.
Frequently Asked Questions
Does the medical tokenizer change how the model sees input text?
No. According to the maziyarpanahi/openmed source code, the tokenizer runs strictly after inference inside openmed/__init__.py. It remaps output spans, confidence scores, and metadata only, leaving the model's sub-word tokenization completely unchanged.
How do I force a multi-word phrase to remain a single token?
Pass the phrase to the exceptions parameter of medical_tokenize, or add it to OpenMedConfig.medical_tokenizer_exceptions. Alternatively, set the OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS environment variable so the phrase is preserved across all pipeline runs.
What is the difference between medical_tokenize and remap_predictions_to_tokens?
medical_tokenize performs the clinical text splitting and returns a list of SpanToken objects. remap_predictions_to_tokens takes the model's raw character-span predictions and aligns them to those tokens, choosing the highest-scoring label per token and merging adjacent tokens that share the same label.
Is there a built-in list of clinical exceptions?
Yes. The constant DEFAULT_MEDICAL_EXCEPTIONS defined at the top of openmed/processing/tokenization.py includes common clinical strings such as COVID-19, CAR-T, and numeric temperature values. Any custom exceptions you provide are merged with this default list.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →