# What Is OpenMed? A Local‑First Healthcare AI Library for Clinical NLP and PII De‑Identification

> Discover OpenMed, the local-first Python library for clinical NLP and PII de-identification. Process sensitive health data securely on-device without external servers.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: getting-started
- Published: 2026-06-13

---

**OpenMed is a Python‑first, open‑source library that executes clinical NLP tasks—including medical named‑entity recognition (NER) and HIPAA‑grade PII de‑identification—entirely on‑device without transmitting data to external servers.**

OpenMed, hosted in the `maziyarpanahi/openmed` repository, provides a modular architecture for transforming unstructured biomedical text into structured data. Designed for healthcare environments with strict privacy requirements, it operates through a **model‑registry → loader → processing pipeline** that supports air‑gapped deployments and multiple acceleration backends.

## Core Capabilities of OpenMed

### Medical Named‑Entity Recognition (NER)

OpenMed identifies clinical entities such as diseases, drugs, anatomical structures, and genes using specialized transformer models. The `analyze_text` function in [`openmed/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/__init__.py) orchestrates sentence segmentation via [`openmed/processing/sentences.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/sentences.py), optional medical‑aware token remapping in [`openmed/processing/tokenization.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/tokenization.py), and entity merging to produce normalized results.

### PII Detection and De‑Identification

The library implements HIPAA Safe Harbor compliance through the [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py) module, which exposes `extract_pii` and `deidentify` functions. These utilities detect all 18 identifiers—including names, dates, SSNs, and medical record numbers—and support masking, replacement, or hashing strategies. Multilingual PII patterns are managed in [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py), enabling accurate redaction across 12 languages.

### Zero‑Shot and Multilingual Extraction

Beyond pretrained medical models, OpenMed supports GLiNER‑style zero‑shot extraction through a unified interface. Language‑specific model mappings in the registry allow seamless switching between English, Spanish, German, and other clinical corpora without code changes.

### Accelerated Local Inference

Performance optimization is handled through multiple backends:
- **Hugging‑Face Transformers** for CPU and CUDA execution
- **MLX** framework for Apple Silicon GPU acceleration
- **CoreML** export for iOS and macOS deployment via OpenMedKit

## Architecture and Source Code Structure

### Model Registry and Manifest

The [`openmed/core/model_registry.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/model_registry.py) file parses a static `models.jsonl` manifest to build the `OPENMED_MODELS` dictionary, mapping model keys to metadata including category, size, and supported languages. This registry enables the `list_models()` and `get_models_by_category()` helper functions.

### Global Configuration

`OpenMedConfig` in [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py) defines a dataclass governing cache directories, device selection (CPU/GPU), logging levels, and environment profiles (`dev`, `prod`, `test`, `fast`). Configuration instances are passed to processing functions to control backend selection and runtime behavior.

### Model Loading Pipeline

The `ModelLoader` class in [`openmed/core/models.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/models.py) constructs Hugging‑Face pipelines with `task="token‑classification"`, handling backend initialization and model weight loading. This component supports local path resolution, allowing air‑gapped installations to load weights from `model_id="./models/..."` without Hub connectivity.

### Processing Layer

The `openmed/processing/` directory contains three critical components:
- **Sentence segmentation** ([`sentences.py`](https://github.com/maziyarpanahi/openmed/blob/main/sentences.py)): Wraps `pySBD` for language‑aware text splitting
- **Tokenization** ([`tokenization.py`](https://github.com/maziyarpanahi/openmed/blob/main/tokenization.py)): Optional medical‑term‑aware token remapping
- **Output formatting** ([`outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/outputs.py)): Implements `PredictionResult` and `format_predictions` for dict, JSON, HTML, and CSV serialization

### PII and Internationalization

[`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py) contains the core de‑identification logic, while [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py) manages language‑specific regex patterns and default model mappings for multilingual PII detection.

### Profiling and Utilities

[`openmed/utils/profiling.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/profiling.py) provides simple CPU‑time profiling utilities used internally by the library to benchmark pipeline performance.

### Example Scripts

The repository includes reference implementations such as [`examples/pii_batch_processing.py`](https://github.com/maziyarpanahi/openmed/blob/main/examples/pii_batch_processing.py), demonstrating batch PII extraction workflows on document collections.

### High‑Level API Entry Points

[`openmed/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/__init__.py) consolidates public functions including `analyze_text`, `extract_pii`, `deidentify`, and `BatchProcessor`. This module also re‑exports `OpenMedConfig` and model discovery utilities, providing a unified import surface for end users.

## Practical Implementation Examples

### Running Medical NER

Extract diseases and medications with confidence scores:

```python
from openmed import analyze_text

result = analyze_text(
    "Patient started on imatinib for chronic myeloid leukemia.",
    model_name="disease_detection_superclinical",
)

for ent in result.entities:
    print(f"{ent.label:<12} {ent.text:<28} {ent.confidence:.2f}")

# DISEASE      chronic myeloid leukemia     0.98

# DRUG         imatinib                     0.95

```

### Detecting and Masking PII

Identify and redact protected health information:

```python
from openmed import extract_pii, deidentify

text = "Patient: John Doe, DOB: 01/15/1970, SSN: 123-45-6789"

# Detect PII (smart merging keeps the full DOB together)

pii = extract_pii(text, model_name="pii_superclinical_large", use_smart_merging=True)

print(pii.entities)   # → list of PIIEntity objects

# Mask all identifiers

masked = deidentify(text, method="mask")
print(masked)          # "[NAME] ... [DATE] ... [SSN]"

```

### Processing Clinical Text in Batches

Handle multiple documents efficiently with the batch processor:

```python
from openmed import BatchProcessor

batch = BatchProcessor(
    model_name="pharma_detection_superclinical",
    group_entities=True,
)

texts = [
    "The patient was prescribed 75 mg clopidogrel for NSTEMI.",
    "Administer 2 mg dexamethasone intravenously."
]

results = batch.process_texts(texts)

for r in results:
    print(r.entities)

```

### Leveraging MLX on Apple Silicon

Install and configure the MLX backend for M1/M2/M3 chips:

```bash
pip install "openmed[mlx]"   # installs the MLX accelerated backend

```

```python
from openmed import analyze_text, OpenMedConfig

cfg = OpenMedConfig(backend="mlx", device="gpu")
result = analyze_text(
    "Patient presents with hypertension.",
    model_name="disease_detection_superclinical",
    config=cfg,
)

```

### Querying the Model Registry

Discover available models by category:

```python
from openmed import list_models, get_models_by_category

print(list_models())                     # all model keys

cardio_models = get_models_by_category("Disease")
print([m.display_name for m in cardio_models])

```

## Summary

- OpenMed processes clinical text through a **local‑first pipeline** defined in `openmed/core/` and `openmed/processing/`, ensuring data never leaves the host machine unless explicitly configured.
- The library provides **medical NER** via `analyze_text` and **HIPAA‑compliant PII handling** via `extract_pii` and `deidentify`, with multilingual support across 12 languages.
- Architecture components include the **Model Registry** ([`openmed/core/model_registry.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/model_registry.py)), **Configuration** ([`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py)), and **Processing Layer** ([`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py)), creating a modular extensible stack.
- Accelerated inference options include Hugging‑Face (CPU/CUDA), **MLX** for Apple Silicon, and CoreML export for iOS/macOS integration through OpenMedKit.

## Frequently Asked Questions

### How does OpenMed ensure HIPAA compliance for clinical text?

OpenMed implements the `deidentify` function in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py) to mask, replace, or hash all 18 Safe Harbor identifiers specified by HIPAA. The [`openmed/core/pii_i18n.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_i18n.py) module provides language‑specific regex patterns and model mappings to ensure accurate redaction across multilingual documents without external API calls.

### Can OpenMed operate in air‑gapped environments without internet access?

Yes. The `ModelLoader` in [`openmed/core/models.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/models.py) supports local path resolution via `model_id="./models/..."`. When models are pre‑downloaded to the local filesystem, the library performs zero network calls, making it suitable for secure hospital networks and offline research environments.

### What is the difference between `analyze_text` and `extract_pii`?

`analyze_text` (exported from [`openmed/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/__init__.py)) performs medical NER to identify clinical entities like diseases, drugs, and anatomical structures. `extract_pii` specifically targets personal identifiers such as names, SSNs, and dates, utilizing smart merging logic in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py) to keep multi‑token entities (like full dates or addresses) intact during extraction.

### How can I accelerate OpenMed inference on Apple Silicon Macs?

Install the optional MLX dependency with `pip install "openmed[mlx]"`, then instantiate `OpenMedConfig(backend="mlx", device="gpu")` from [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py). Pass this configuration to `analyze_text` or `BatchProcessor` to leverage the unified memory architecture of M1/M2/M3 chips for faster local inference.