# How OpenMed's Medical Tokenizer Remaps Predictions onto Stable Medical Tokens

> OpenMed's medical tokenizer remaps predictions onto stable tokens by splitting text, aligning predictions, and merging adjacent tokens. Learn how this process works.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: internals
- Published: 2026-06-10

---

**OpenMed remaps character-based model predictions onto stable medical tokens by first splitting text into immutable `SpanToken` objects, then aligning predictions to the highest-scoring overlapping tokens, and finally merging adjacent tokens that share the same entity label within a configurable character gap.**

The `maziyarpanahi/openmed` repository provides a medical-aware tokenizer that operates exclusively during output post-processing to transform raw model predictions into human-readable medical entities. Unlike input tokenizers that affect how models receive text, this component remaps sub-word or character-level predictions onto stable clinical token boundaries. Understanding how the OpenMed medical tokenizer remaps predictions is essential for interpreting named entity recognition results in clinical workflows.

## The Three-Step Remapping Pipeline

### Step 1: Generating Stable Medical Spans

The process begins in [`openmed/processing/tokenization.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/tokenization.py) with the `medical_tokenize` function, which applies the `_MEDICAL_TOKEN_PATTERN` regular expression to split clinical strings into immutable `SpanToken` objects containing `text`, `start`, and `end` attributes. This pattern recognizes numbers, hyphenated words, ratios, and protected exception terms like "COVID-19" that should not be split, returning a sorted list of spans that serve as the anchor points for remapping.

### Step 2: Aligning Predictions to Token Spans

The `remap_predictions_to_tokens` function receives raw predictions—character-based start/end spans with labels, scores, and optional metadata—alongside the original text and the token list generated in step one. It iterates over all predictions to build per-token mappings, keeping only the highest-scoring overlapping label for each `SpanToken` according to the logic in lines 107-135 of the tokenization module.

### Step 3: Merging Adjacent Medical Entities

After per-token labeling, the function walks the token list to concatenate consecutive tokens sharing the same label where the distance between start positions falls within the `gap` parameter (defaulting to 1 character). The merged entity receives the first token's start position, the last token's end position, an averaged score, the cleaned label with B-/I- prefixes stripped, the original text slice, and metadata from the first contributing token, as implemented in lines 137-166.

## Configuration and Integration

According to [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py), the remapping step activates when the `medical_tokenizer` configuration flag is enabled. The high-level inference pipeline wires this functionality through [`openmed/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/__init__.py) (lines 473-486), ensuring that model outputs undergo medical tokenization before final presentation without altering the underlying model's input processing.

## Practical Implementation Example

The following example demonstrates how sub-word predictions covering "IL-6-mediated" are remapped onto a single stable medical token:

```python
from openmed.processing.tokenization import medical_tokenize, remap_predictions_to_tokens

text = "IL-6-mediated cytokine storm"

# 1️⃣ Tokenize into stable medical spans (protects hyphen‑chains, ratios, etc.)

tokens = medical_tokenize(text)

# tokens → [SpanToken(text='IL-6-mediated', start=0, end=13), …]

# 2️⃣ Simulated model predictions on sub‑word pieces (character spans)

preds = [
    {"start": 0, "end": 2, "entity": "B-Gene_or_gene_product", "score": 0.9, "metadata": {"sentence_index": 0}},
    {"start": 3, "end": 4, "entity": "I-Gene_or_gene_product", "score": 0.8, "metadata": {"sentence_index": 0}},
    {"start": 5, "end": 13, "entity": "I-Gene_or_gene_product", "score": 0.85, "metadata": {"sentence_index": 0}},
]

# 3️⃣ Remap predictions onto the stable medical tokens

remapped = remap_predictions_to_tokens(preds, text, tokens)

# remapped → [{

#   "start": 0,

#   "end": 13,

#   "score": 0.85,

#   "entity_group": "Gene_or_gene_product",

#   "word": "IL-6-mediated",

#   "metadata": {"sentence_index": 0}

# }]

print(remapped)

```

## Validating the Remapping Logic

The [`tests/test_medical_remap.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/test_medical_remap.py) file validates this behavior through unit tests including `test_remap_predictions_merges_wordpieces_to_medical_token`, which confirms that three separate B-I-I predictions covering "IL-6-mediated" merge into a single span with the cleaned label *Gene_or_gene_product*. Additional tests verify the `gap` parameter functionality for merging adjacent tokens when characters separate them.

## Summary

- OpenMed's medical tokenizer operates only during output post-processing, leaving model input tokenization unchanged.
- The `medical_tokenize` function in [`openmed/processing/tokenization.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/tokenization.py) creates immutable `SpanToken` objects using regex patterns that protect medical terms like "COVID-19".
- `remap_predictions_to_tokens` aligns character-based predictions to tokens by selecting the highest-scoring overlapping label for each span.
- Adjacent tokens with identical labels are merged when their start positions fall within the configurable `gap` distance, producing final entities with averaged scores and cleaned labels.

## Frequently Asked Questions

### Does the medical tokenizer change how the model receives input?

No. According to the source code in `maziyarpanahi/openmed`, the medical tokenizer is used exclusively for output post-processing. It does not affect how the model receives its input IDs or the internal tokenization used during inference.

### How does OpenMed handle overlapping predictions during remapping?

When multiple predictions overlap a single token span, the `remap_predictions_to_tokens` function keeps only the highest-scoring label for that span. This per-token best-label selection occurs in lines 107-135 of [`openmed/processing/tokenization.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/tokenization.py).

### What happens to BIO tags during the remapping process?

The remapping logic strips B- and I- prefixes from entity labels during the merge phase. The final output uses the cleaned `entity_group` field without these prefixes, representing the consolidated medical entity type.

### Can I adjust how aggressively the tokenizer merges adjacent tokens?

Yes. The `remap_predictions_to_tokens` function accepts a `gap` parameter (defaulting to 1) that determines how many characters may separate adjacent tokens while still allowing them to merge into a single entity. Adjust this parameter to control boundary detection strictness.