How OpenMed's Medical Tokenizer Remaps Predictions onto Stable Medical Tokens
OpenMed remaps character-based model predictions onto stable medical tokens by first splitting text into immutable SpanToken objects, then aligning predictions to the highest-scoring overlapping tokens, and finally merging adjacent tokens that share the same entity label within a configurable character gap.
The maziyarpanahi/openmed repository provides a medical-aware tokenizer that operates exclusively during output post-processing to transform raw model predictions into human-readable medical entities. Unlike input tokenizers that affect how models receive text, this component remaps sub-word or character-level predictions onto stable clinical token boundaries. Understanding how the OpenMed medical tokenizer remaps predictions is essential for interpreting named entity recognition results in clinical workflows.
The Three-Step Remapping Pipeline
Step 1: Generating Stable Medical Spans
The process begins in openmed/processing/tokenization.py with the medical_tokenize function, which applies the _MEDICAL_TOKEN_PATTERN regular expression to split clinical strings into immutable SpanToken objects containing text, start, and end attributes. This pattern recognizes numbers, hyphenated words, ratios, and protected exception terms like "COVID-19" that should not be split, returning a sorted list of spans that serve as the anchor points for remapping.
Step 2: Aligning Predictions to Token Spans
The remap_predictions_to_tokens function receives raw predictions—character-based start/end spans with labels, scores, and optional metadata—alongside the original text and the token list generated in step one. It iterates over all predictions to build per-token mappings, keeping only the highest-scoring overlapping label for each SpanToken according to the logic in lines 107-135 of the tokenization module.
Step 3: Merging Adjacent Medical Entities
After per-token labeling, the function walks the token list to concatenate consecutive tokens sharing the same label where the distance between start positions falls within the gap parameter (defaulting to 1 character). The merged entity receives the first token's start position, the last token's end position, an averaged score, the cleaned label with B-/I- prefixes stripped, the original text slice, and metadata from the first contributing token, as implemented in lines 137-166.
Configuration and Integration
According to openmed/core/config.py, the remapping step activates when the medical_tokenizer configuration flag is enabled. The high-level inference pipeline wires this functionality through openmed/__init__.py (lines 473-486), ensuring that model outputs undergo medical tokenization before final presentation without altering the underlying model's input processing.
Practical Implementation Example
The following example demonstrates how sub-word predictions covering "IL-6-mediated" are remapped onto a single stable medical token:
from openmed.processing.tokenization import medical_tokenize, remap_predictions_to_tokens
text = "IL-6-mediated cytokine storm"
# 1️⃣ Tokenize into stable medical spans (protects hyphen‑chains, ratios, etc.)
tokens = medical_tokenize(text)
# tokens → [SpanToken(text='IL-6-mediated', start=0, end=13), …]
# 2️⃣ Simulated model predictions on sub‑word pieces (character spans)
preds = [
{"start": 0, "end": 2, "entity": "B-Gene_or_gene_product", "score": 0.9, "metadata": {"sentence_index": 0}},
{"start": 3, "end": 4, "entity": "I-Gene_or_gene_product", "score": 0.8, "metadata": {"sentence_index": 0}},
{"start": 5, "end": 13, "entity": "I-Gene_or_gene_product", "score": 0.85, "metadata": {"sentence_index": 0}},
]
# 3️⃣ Remap predictions onto the stable medical tokens
remapped = remap_predictions_to_tokens(preds, text, tokens)
# remapped → [{
# "start": 0,
# "end": 13,
# "score": 0.85,
# "entity_group": "Gene_or_gene_product",
# "word": "IL-6-mediated",
# "metadata": {"sentence_index": 0}
# }]
print(remapped)
Validating the Remapping Logic
The tests/test_medical_remap.py file validates this behavior through unit tests including test_remap_predictions_merges_wordpieces_to_medical_token, which confirms that three separate B-I-I predictions covering "IL-6-mediated" merge into a single span with the cleaned label Gene_or_gene_product. Additional tests verify the gap parameter functionality for merging adjacent tokens when characters separate them.
Summary
- OpenMed's medical tokenizer operates only during output post-processing, leaving model input tokenization unchanged.
- The
medical_tokenizefunction inopenmed/processing/tokenization.pycreates immutableSpanTokenobjects using regex patterns that protect medical terms like "COVID-19". remap_predictions_to_tokensaligns character-based predictions to tokens by selecting the highest-scoring overlapping label for each span.- Adjacent tokens with identical labels are merged when their start positions fall within the configurable
gapdistance, producing final entities with averaged scores and cleaned labels.
Frequently Asked Questions
Does the medical tokenizer change how the model receives input?
No. According to the source code in maziyarpanahi/openmed, the medical tokenizer is used exclusively for output post-processing. It does not affect how the model receives its input IDs or the internal tokenization used during inference.
How does OpenMed handle overlapping predictions during remapping?
When multiple predictions overlap a single token span, the remap_predictions_to_tokens function keeps only the highest-scoring label for that span. This per-token best-label selection occurs in lines 107-135 of openmed/processing/tokenization.py.
What happens to BIO tags during the remapping process?
The remapping logic strips B- and I- prefixes from entity labels during the merge phase. The final output uses the cleaned entity_group field without these prefixes, representing the consolidated medical entity type.
Can I adjust how aggressively the tokenizer merges adjacent tokens?
Yes. The remap_predictions_to_tokens function accepts a gap parameter (defaulting to 1) that determines how many characters may separate adjacent tokens while still allowing them to merge into a single entity. Adjust this parameter to control boundary detection strictness.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →