What Is OpenMed? A Local‑First Healthcare AI Library for Clinical NLP and PII De‑Identification
OpenMed is a Python‑first, open‑source library that executes clinical NLP tasks—including medical named‑entity recognition (NER) and HIPAA‑grade PII de‑identification—entirely on‑device without transmitting data to external servers.
OpenMed, hosted in the maziyarpanahi/openmed repository, provides a modular architecture for transforming unstructured biomedical text into structured data. Designed for healthcare environments with strict privacy requirements, it operates through a model‑registry → loader → processing pipeline that supports air‑gapped deployments and multiple acceleration backends.
Core Capabilities of OpenMed
Medical Named‑Entity Recognition (NER)
OpenMed identifies clinical entities such as diseases, drugs, anatomical structures, and genes using specialized transformer models. The analyze_text function in openmed/__init__.py orchestrates sentence segmentation via openmed/processing/sentences.py, optional medical‑aware token remapping in openmed/processing/tokenization.py, and entity merging to produce normalized results.
PII Detection and De‑Identification
The library implements HIPAA Safe Harbor compliance through the openmed/core/pii.py module, which exposes extract_pii and deidentify functions. These utilities detect all 18 identifiers—including names, dates, SSNs, and medical record numbers—and support masking, replacement, or hashing strategies. Multilingual PII patterns are managed in openmed/core/pii_i18n.py, enabling accurate redaction across 12 languages.
Zero‑Shot and Multilingual Extraction
Beyond pretrained medical models, OpenMed supports GLiNER‑style zero‑shot extraction through a unified interface. Language‑specific model mappings in the registry allow seamless switching between English, Spanish, German, and other clinical corpora without code changes.
Accelerated Local Inference
Performance optimization is handled through multiple backends:
- Hugging‑Face Transformers for CPU and CUDA execution
- MLX framework for Apple Silicon GPU acceleration
- CoreML export for iOS and macOS deployment via OpenMedKit
Architecture and Source Code Structure
Model Registry and Manifest
The openmed/core/model_registry.py file parses a static models.jsonl manifest to build the OPENMED_MODELS dictionary, mapping model keys to metadata including category, size, and supported languages. This registry enables the list_models() and get_models_by_category() helper functions.
Global Configuration
OpenMedConfig in openmed/core/config.py defines a dataclass governing cache directories, device selection (CPU/GPU), logging levels, and environment profiles (dev, prod, test, fast). Configuration instances are passed to processing functions to control backend selection and runtime behavior.
Model Loading Pipeline
The ModelLoader class in openmed/core/models.py constructs Hugging‑Face pipelines with task="token‑classification", handling backend initialization and model weight loading. This component supports local path resolution, allowing air‑gapped installations to load weights from model_id="./models/..." without Hub connectivity.
Processing Layer
The openmed/processing/ directory contains three critical components:
- Sentence segmentation (
sentences.py): WrapspySBDfor language‑aware text splitting - Tokenization (
tokenization.py): Optional medical‑term‑aware token remapping - Output formatting (
outputs.py): ImplementsPredictionResultandformat_predictionsfor dict, JSON, HTML, and CSV serialization
PII and Internationalization
openmed/core/pii.py contains the core de‑identification logic, while openmed/core/pii_i18n.py manages language‑specific regex patterns and default model mappings for multilingual PII detection.
Profiling and Utilities
openmed/utils/profiling.py provides simple CPU‑time profiling utilities used internally by the library to benchmark pipeline performance.
Example Scripts
The repository includes reference implementations such as examples/pii_batch_processing.py, demonstrating batch PII extraction workflows on document collections.
High‑Level API Entry Points
openmed/__init__.py consolidates public functions including analyze_text, extract_pii, deidentify, and BatchProcessor. This module also re‑exports OpenMedConfig and model discovery utilities, providing a unified import surface for end users.
Practical Implementation Examples
Running Medical NER
Extract diseases and medications with confidence scores:
from openmed import analyze_text
result = analyze_text(
"Patient started on imatinib for chronic myeloid leukemia.",
model_name="disease_detection_superclinical",
)
for ent in result.entities:
print(f"{ent.label:<12} {ent.text:<28} {ent.confidence:.2f}")
# DISEASE chronic myeloid leukemia 0.98
# DRUG imatinib 0.95
Detecting and Masking PII
Identify and redact protected health information:
from openmed import extract_pii, deidentify
text = "Patient: John Doe, DOB: 01/15/1970, SSN: 123-45-6789"
# Detect PII (smart merging keeps the full DOB together)
pii = extract_pii(text, model_name="pii_superclinical_large", use_smart_merging=True)
print(pii.entities) # → list of PIIEntity objects
# Mask all identifiers
masked = deidentify(text, method="mask")
print(masked) # "[NAME] ... [DATE] ... [SSN]"
Processing Clinical Text in Batches
Handle multiple documents efficiently with the batch processor:
from openmed import BatchProcessor
batch = BatchProcessor(
model_name="pharma_detection_superclinical",
group_entities=True,
)
texts = [
"The patient was prescribed 75 mg clopidogrel for NSTEMI.",
"Administer 2 mg dexamethasone intravenously."
]
results = batch.process_texts(texts)
for r in results:
print(r.entities)
Leveraging MLX on Apple Silicon
Install and configure the MLX backend for M1/M2/M3 chips:
pip install "openmed[mlx]" # installs the MLX accelerated backend
from openmed import analyze_text, OpenMedConfig
cfg = OpenMedConfig(backend="mlx", device="gpu")
result = analyze_text(
"Patient presents with hypertension.",
model_name="disease_detection_superclinical",
config=cfg,
)
Querying the Model Registry
Discover available models by category:
from openmed import list_models, get_models_by_category
print(list_models()) # all model keys
cardio_models = get_models_by_category("Disease")
print([m.display_name for m in cardio_models])
Summary
- OpenMed processes clinical text through a local‑first pipeline defined in
openmed/core/andopenmed/processing/, ensuring data never leaves the host machine unless explicitly configured. - The library provides medical NER via
analyze_textand HIPAA‑compliant PII handling viaextract_piianddeidentify, with multilingual support across 12 languages. - Architecture components include the Model Registry (
openmed/core/model_registry.py), Configuration (openmed/core/config.py), and Processing Layer (openmed/processing/outputs.py), creating a modular extensible stack. - Accelerated inference options include Hugging‑Face (CPU/CUDA), MLX for Apple Silicon, and CoreML export for iOS/macOS integration through OpenMedKit.
Frequently Asked Questions
How does OpenMed ensure HIPAA compliance for clinical text?
OpenMed implements the deidentify function in openmed/core/pii.py to mask, replace, or hash all 18 Safe Harbor identifiers specified by HIPAA. The openmed/core/pii_i18n.py module provides language‑specific regex patterns and model mappings to ensure accurate redaction across multilingual documents without external API calls.
Can OpenMed operate in air‑gapped environments without internet access?
Yes. The ModelLoader in openmed/core/models.py supports local path resolution via model_id="./models/...". When models are pre‑downloaded to the local filesystem, the library performs zero network calls, making it suitable for secure hospital networks and offline research environments.
What is the difference between analyze_text and extract_pii?
analyze_text (exported from openmed/__init__.py) performs medical NER to identify clinical entities like diseases, drugs, and anatomical structures. extract_pii specifically targets personal identifiers such as names, SSNs, and dates, utilizing smart merging logic in openmed/core/pii.py to keep multi‑token entities (like full dates or addresses) intact during extraction.
How can I accelerate OpenMed inference on Apple Silicon Macs?
Install the optional MLX dependency with pip install "openmed[mlx]", then instantiate OpenMedConfig(backend="mlx", device="gpu") from openmed/core/config.py. Pass this configuration to analyze_text or BatchProcessor to leverage the unified memory architecture of M1/M2/M3 chips for faster local inference.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →