# Viterbi Decoding Strategy for NER in OpenMed: BIOES-Aware Implementation Guide

> Learn the Viterbi decoding strategy for NER in OpenMed. This guide details the BIOES-aware implementation, transforming log-probabilities into valid named entity spans using dynamic programming.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-10

---

**OpenMed employs a BIOES-aware Viterbi decoder in [`openmed/core/decoding/viterbi.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/decoding/viterbi.py) to transform raw token classification log-probabilities into valid named entity spans by enforcing structural constraints through dynamic programming.**

The `maziyarpanahi/openmed` repository implements a privacy-focused medical NER system that relies on structured prediction to ensure output validity. At the core of this approach lies a pure-Python Viterbi decoding strategy that operates on BIOES-tagged sequences without requiring heavy tensor libraries, making it callable from Swift, native Python, or other runtimes.

## How the BIOES-Aware Viterbi Decoder Works in OpenMed

Unlike simple argmax decoding, OpenMed’s strategy enforces valid **BIOES** (Begin, Inside, Outside, End, Singleton) transitions at inference time. The decoder accepts raw log-probabilities from privacy-filter models and applies constrained dynamic programming to eliminate impossible sequences (such as an `I-NAME` token following an `O` token).

### Core Architecture and File Location

The complete implementation resides in [`openmed/core/decoding/viterbi.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/decoding/viterbi.py). This module exposes `viterbi_decode()`, `build_label_info()`, and supporting utilities that construct transition matrices, apply learned biases, and recover optimal label paths. The decoder is backend-agnostic, operating on standard Python `list[list[float]]` structures, which allows identical usage across both MLX (Apple Silicon) and PyTorch inference pipelines.

## The Six-Step Viterbi Decoding Pipeline

### 1. Building Label Metadata with `build_label_info()`

The decoder first constructs a `TokenLabelInfo` object via `build_label_info(id2label)` (lines 74-78). This preprocessing step:

- Extracts span-label indices from the model’s `id2label` mapping (e.g., `0 → "O"`, `1 → "B-NAME"`)
- Maps each token label to its BIOES boundary tag (`B`, `I`, `E`, `S`, or background)
- Pre-computes valid transition lookups to accelerate later scoring

### 2. Preparing Transition Biases

OpenMed supports six optional learned scalar biases that adjust transition probabilities. The function `zero_viterbi_biases()` (lines 80-83) initializes a default bias dictionary with zero values. When provided, user-specified biases override these defaults (lines 37-39), allowing the model to learn preferences for specific state transitions such as *background → start* or *inside → continue*.

### 3. Constructing Viterbi Score Matrices

The internal `_build_viterbi_scores()` function (lines 66-108) constructs three constrained score structures:

- **Start scores**: Permitted only for `B`, `S`, or background tokens
- **End scores**: Permitted only for `E`, `S`, or background tokens  
- **Transition scores**: Populated exclusively for valid BIOES transitions as validated by `_is_valid_transition()`, then weighted by `_transition_bias()`

### 4. Dynamic Programming Optimization

The `viterbi_decode()` function executes the classic Viterbi recursion (lines 50-65). For each token position, it computes the maximum score path by combining:

1. The previous token’s accumulated scores
2. The current transition score
3. The current emission score (from input logits)

Back-pointers are stored at each step to enable path reconstruction.

### 5. Path Extraction and Termination

Upon reaching the sequence end, final scores incorporate end-score constraints. The algorithm selects the highest-scoring final label, then traverses the stored back-pointers backward (lines 66-76) to recover the complete optimal label sequence that respects BIOES boundaries.

### 6. Converting Labels to NER Spans

The `labels_to_token_spans()` function (lines 78-115) transforms the decoded label indices into `(span_label, start, end)` triples. This utility respects BIOES semantics: singleton tags (`S`) create single-token spans, while `B-I-E` sequences merge into multi-token entities. Background `O` tokens automatically terminate open spans, and gaps in token indices break entity continuity.

## Implementation Example: Decoding Token Logits

The following example mirrors the unit tests in [`tests/unit/mlx/test_privacy_filter_mlx.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/mlx/test_privacy_filter_mlx.py) and demonstrates the complete decoding workflow:

```python
from openmed.core.decoding import build_label_info, viterbi_decode

# 1. Model's id-to-label map (example)

id2label = {
    0: "O",
    1: "B-NAME",
    2: "I-NAME",
    3: "E-NAME",
    4: "S-EMAIL",
}

# 2. Build metadata

label_info = build_label_info(id2label)

# 3. Fake token log-probabilities (log-softmax values)

#    shape: [num_tokens, num_classes]

token_logprobs = [
    [-1.2, -0.3, -2.0, -2.5, -3.0],  # token 0

    [-0.5, -0.1, -0.2, -3.0, -2.0],  # token 1

    [-0.7, -2.5, -0.1, -0.2, -3.0],  # token 2

]

# 4. (optional) transition biases learned by the model

biases = {
    "transition_bias_background_to_start": 0.5,
    "transition_bias_inside_to_continue": 0.2,
    # other bias keys default to 0.0

}

# 5. Decode the best BIOES-valid label path

predicted_ids = viterbi_decode(token_logprobs, label_info=label_info, biases=biases)

print(predicted_ids)   # e.g. [1, 2, 3] → B-NAME, I-NAME, E-NAME

```

This same pattern executes inside OpenMed’s inference pipelines; developers supply only the logits and `id2label` mapping to obtain structurally valid predictions.

## Integration with MLX and PyTorch Inference

The decoder integrates directly into OpenMed’s inference stack. In [`openmed/mlx/inference.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/inference.py) at line 482, the MLX inference module calls `viterbi_decode()` immediately after obtaining per-token log-probabilities from the privacy-filter model:

```python
pred_ids = viterbi_decode(logits, label_info=label_info, biases=model_biases)

```

The PyTorch wrapper reuses identical functions through the public API exported in [`openmed/core/decoding/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/decoding/__init__.py). Because the decoder is model-led—incorporating learned transition biases—it eliminates the need for separate rule-based post-processors while ensuring 100% BIOES-compliant outputs.

## Summary

- **Location**: The Viterbi decoder is implemented in [`openmed/core/decoding/viterbi.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/decoding/viterbi.py) with public exports in [`openmed/core/decoding/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/decoding/__init__.py).
- **Method**: BIOES-aware dynamic programming enforces valid entity boundaries during decoding rather than filtering invalid sequences afterward.
- **Inputs**: Raw token log-probabilities, an `id2label` mapping, and optional learned transition biases.
- **Outputs**: Valid BIOES label sequences converted to `(span_label, start, end)` tuples via `labels_to_token_spans()`.
- **Portability**: Pure-Python implementation with no tensor dependencies enables deployment across MLX, PyTorch, Swift wrappers, and other runtimes.

## Frequently Asked Questions

### What makes OpenMed's Viterbi decoder "BIOES-aware"?

The decoder explicitly encodes BIOES constraints into its transition matrix. The `_is_valid_transition()` function validates that sequences follow proper Begin-Inside-End-Singleton semantics, preventing illegal transitions like `I-NAME` following `O` or `B-NAME` following `I-NAME`. This structural awareness is baked into the `_build_viterbi_scores()` logic rather than applied as a post-processing filter.

### How does the decoder integrate with MLX inference pipelines?

In [`openmed/mlx/inference.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/inference.py) at line 482, the MLX inference engine passes model outputs directly to `viterbi_decode()` after computing log-softmax probabilities. The decoder returns integer label IDs that map to entity spans, allowing the privacy filter to redact sensitive medical information without additional validation steps.

### What are transition biases and how do they affect decoding?

Transition biases are six optional scalar values learned during model training that adjust the Viterbi transition matrix. Stored in dictionaries returned by `zero_viterbi_biases()`, these parameters allow the model to encode preferences such as favoring the continuation of an entity (`inside_to_continue`) or biasing against invalid background-to-inside jumps. When provided, these biases are added to transition scores via `_transition_bias()` before dynamic programming begins.

### Can the Viterbi decoder be used outside of OpenMed?

Yes. Because the implementation in [`openmed/core/decoding/viterbi.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/decoding/viterbi.py) relies only on Python standard library features and operates on `list[list[float]]` inputs, it can decode outputs from any token-classification model that produces log-probability matrices. The decoder requires only an `id2label` mapping and optionally a bias dictionary, making it portable to custom NER pipelines beyond the OpenMed ecosystem.