What Is OpenMed? A Local‑First Healthcare AI Library for Clinical NLP and PII De‑Identification

OpenMed is a Python‑first, open‑source library that executes clinical NLP tasks—including medical named‑entity recognition (NER) and HIPAA‑grade PII de‑identification—entirely on‑device without transmitting data to external servers.

OpenMed, hosted in the maziyarpanahi/openmed repository, provides a modular architecture for transforming unstructured biomedical text into structured data. Designed for healthcare environments with strict privacy requirements, it operates through a model‑registry → loader → processing pipeline that supports air‑gapped deployments and multiple acceleration backends.

Core Capabilities of OpenMed

Medical Named‑Entity Recognition (NER)

OpenMed identifies clinical entities such as diseases, drugs, anatomical structures, and genes using specialized transformer models. The analyze_text function in openmed/__init__.py orchestrates sentence segmentation via openmed/processing/sentences.py, optional medical‑aware token remapping in openmed/processing/tokenization.py, and entity merging to produce normalized results.

PII Detection and De‑Identification

The library implements HIPAA Safe Harbor compliance through the openmed/core/pii.py module, which exposes extract_pii and deidentify functions. These utilities detect all 18 identifiers—including names, dates, SSNs, and medical record numbers—and support masking, replacement, or hashing strategies. Multilingual PII patterns are managed in openmed/core/pii_i18n.py, enabling accurate redaction across 12 languages.

Zero‑Shot and Multilingual Extraction

Beyond pretrained medical models, OpenMed supports GLiNER‑style zero‑shot extraction through a unified interface. Language‑specific model mappings in the registry allow seamless switching between English, Spanish, German, and other clinical corpora without code changes.

Accelerated Local Inference

Performance optimization is handled through multiple backends:

  • Hugging‑Face Transformers for CPU and CUDA execution
  • MLX framework for Apple Silicon GPU acceleration
  • CoreML export for iOS and macOS deployment via OpenMedKit

Architecture and Source Code Structure

Model Registry and Manifest

The openmed/core/model_registry.py file parses a static models.jsonl manifest to build the OPENMED_MODELS dictionary, mapping model keys to metadata including category, size, and supported languages. This registry enables the list_models() and get_models_by_category() helper functions.

Global Configuration

OpenMedConfig in openmed/core/config.py defines a dataclass governing cache directories, device selection (CPU/GPU), logging levels, and environment profiles (dev, prod, test, fast). Configuration instances are passed to processing functions to control backend selection and runtime behavior.

Model Loading Pipeline

The ModelLoader class in openmed/core/models.py constructs Hugging‑Face pipelines with task="token‑classification", handling backend initialization and model weight loading. This component supports local path resolution, allowing air‑gapped installations to load weights from model_id="./models/..." without Hub connectivity.

Processing Layer

The openmed/processing/ directory contains three critical components:

  • Sentence segmentation (sentences.py): Wraps pySBD for language‑aware text splitting
  • Tokenization (tokenization.py): Optional medical‑term‑aware token remapping
  • Output formatting (outputs.py): Implements PredictionResult and format_predictions for dict, JSON, HTML, and CSV serialization

PII and Internationalization

openmed/core/pii.py contains the core de‑identification logic, while openmed/core/pii_i18n.py manages language‑specific regex patterns and default model mappings for multilingual PII detection.

Profiling and Utilities

openmed/utils/profiling.py provides simple CPU‑time profiling utilities used internally by the library to benchmark pipeline performance.

Example Scripts

The repository includes reference implementations such as examples/pii_batch_processing.py, demonstrating batch PII extraction workflows on document collections.

High‑Level API Entry Points

openmed/__init__.py consolidates public functions including analyze_text, extract_pii, deidentify, and BatchProcessor. This module also re‑exports OpenMedConfig and model discovery utilities, providing a unified import surface for end users.

Practical Implementation Examples

Running Medical NER

Extract diseases and medications with confidence scores:

from openmed import analyze_text

result = analyze_text(
    "Patient started on imatinib for chronic myeloid leukemia.",
    model_name="disease_detection_superclinical",
)

for ent in result.entities:
    print(f"{ent.label:<12} {ent.text:<28} {ent.confidence:.2f}")

# DISEASE      chronic myeloid leukemia     0.98

# DRUG         imatinib                     0.95

Detecting and Masking PII

Identify and redact protected health information:

from openmed import extract_pii, deidentify

text = "Patient: John Doe, DOB: 01/15/1970, SSN: 123-45-6789"

# Detect PII (smart merging keeps the full DOB together)

pii = extract_pii(text, model_name="pii_superclinical_large", use_smart_merging=True)

print(pii.entities)   # → list of PIIEntity objects

# Mask all identifiers

masked = deidentify(text, method="mask")
print(masked)          # "[NAME] ... [DATE] ... [SSN]"

Processing Clinical Text in Batches

Handle multiple documents efficiently with the batch processor:

from openmed import BatchProcessor

batch = BatchProcessor(
    model_name="pharma_detection_superclinical",
    group_entities=True,
)

texts = [
    "The patient was prescribed 75 mg clopidogrel for NSTEMI.",
    "Administer 2 mg dexamethasone intravenously."
]

results = batch.process_texts(texts)

for r in results:
    print(r.entities)

Leveraging MLX on Apple Silicon

Install and configure the MLX backend for M1/M2/M3 chips:

pip install "openmed[mlx]"   # installs the MLX accelerated backend
from openmed import analyze_text, OpenMedConfig

cfg = OpenMedConfig(backend="mlx", device="gpu")
result = analyze_text(
    "Patient presents with hypertension.",
    model_name="disease_detection_superclinical",
    config=cfg,
)

Querying the Model Registry

Discover available models by category:

from openmed import list_models, get_models_by_category

print(list_models())                     # all model keys

cardio_models = get_models_by_category("Disease")
print([m.display_name for m in cardio_models])

Summary

  • OpenMed processes clinical text through a local‑first pipeline defined in openmed/core/ and openmed/processing/, ensuring data never leaves the host machine unless explicitly configured.
  • The library provides medical NER via analyze_text and HIPAA‑compliant PII handling via extract_pii and deidentify, with multilingual support across 12 languages.
  • Architecture components include the Model Registry (openmed/core/model_registry.py), Configuration (openmed/core/config.py), and Processing Layer (openmed/processing/outputs.py), creating a modular extensible stack.
  • Accelerated inference options include Hugging‑Face (CPU/CUDA), MLX for Apple Silicon, and CoreML export for iOS/macOS integration through OpenMedKit.

Frequently Asked Questions

How does OpenMed ensure HIPAA compliance for clinical text?

OpenMed implements the deidentify function in openmed/core/pii.py to mask, replace, or hash all 18 Safe Harbor identifiers specified by HIPAA. The openmed/core/pii_i18n.py module provides language‑specific regex patterns and model mappings to ensure accurate redaction across multilingual documents without external API calls.

Can OpenMed operate in air‑gapped environments without internet access?

Yes. The ModelLoader in openmed/core/models.py supports local path resolution via model_id="./models/...". When models are pre‑downloaded to the local filesystem, the library performs zero network calls, making it suitable for secure hospital networks and offline research environments.

What is the difference between analyze_text and extract_pii?

analyze_text (exported from openmed/__init__.py) performs medical NER to identify clinical entities like diseases, drugs, and anatomical structures. extract_pii specifically targets personal identifiers such as names, SSNs, and dates, utilizing smart merging logic in openmed/core/pii.py to keep multi‑token entities (like full dates or addresses) intact during extraction.

How can I accelerate OpenMed inference on Apple Silicon Macs?

Install the optional MLX dependency with pip install "openmed[mlx]", then instantiate OpenMedConfig(backend="mlx", device="gpu") from openmed/core/config.py. Pass this configuration to analyze_text or BatchProcessor to leverage the unified memory architecture of M1/M2/M3 chips for faster local inference.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →