How to Deploy OpenMed: Complete Guide to Python Package, REST API, and Docker Installation

Deploy OpenMed by installing the Python package with pip install "openmed[hf]", running the FastAPI REST service with uvicorn openmed.service.app:app, or building the Docker image from the repository's Dockerfile to serve the local-first healthcare AI on any infrastructure.

OpenMed is a local-first healthcare AI library developed by maziyarpanahi/openmed that runs completely on-device, ensuring patient data never leaves the host. The repository provides three deployment surfaces: a Python package for direct library integration, a FastAPI REST service defined in openmed/service/app.py, and a production-ready Docker container. This guide covers each deployment method using the actual source code structure and configuration options found in the repository.

Install OpenMed as a Python Package

The quickest way to deploy OpenMed is via pip or uv. The library exposes its public API through openmed/__init__.py, which includes functions like analyze_text, extract_pii, and deidentify.

Install the core library with Hugging Face extras for CPU and CUDA support:

pip install "openmed[hf]"

For Apple Silicon acceleration using the MLX backend, reference the backend selection logic in openmed/core/backends.py and install the MLX variant:

pip install "openmed[mlx]"

After installation, you can call the high-level API directly:

from openmed import analyze_text

result = analyze_text(
    "Patient started on imatinib for chronic myeloid leukemia.",
    model_name="disease_detection_superclinical"
)

print(result.entities)

Run the OpenMed REST Service Locally

The REST service is a FastAPI application defined in openmed/service/app.py that exposes HTTP endpoints for /analyze, /pii/extract, and /pii/deidentify. The service uses a shared ServiceRuntime instantiated in openmed/service/runtime.py to manage model loading and keep-alive policies across requests.

Install Service Dependencies

Install the optional service extra to include FastAPI, uvicorn, and Pydantic models from openmed/service/schemas.py:

pip install "openmed[hf,service]"

Start the Server

Launch the service using uvicorn with the app factory defined in openmed/service/app.py:

uvicorn openmed.service.app:app --host 0.0.0.0 --port 8080

Configure Environment Variables

The runtime behavior is controlled through environment variables read by the ServiceRuntime class. Set these before starting the server:

  • OPENMED_PROFILE: Selects runtime profile (dev, prod, test) affecting request timeouts
  • OPENMED_SERVICE_KEEP_ALIVE: Model cache TTL (e.g., 10m); set to 0 to disable caching
  • OPENMED_SERVICE_PRELOAD_MODELS: Comma-separated list of models to load at startup (e.g., disease_detection_superclinical,OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1)
  • OPENMED_SERVICE_MAX_TEXT_LENGTH: Maximum request payload size for pre-validation (default 250000)

Example request using curl:

curl -X POST http://127.0.0.1:8080/analyze \
  -H "Content-Type: application/json" \
  -d '{"text":"Patient started imatinib for CML.", "model_name":"disease_detection_superclinical"}'

Deploy OpenMed with Docker

For production deployments, the repository includes a Dockerfile that builds a minimal Python-slim image with the service pre-installed.

Build the Image

docker build -t openmed:latest .

Run the Container

The container automatically starts uvicorn openmed.service.app:app via the CMD instruction in the Dockerfile. Pass environment variables to configure the runtime:

docker run --rm -p 8080:8080 \
  -e OPENMED_PROFILE=prod \
  -e OPENMED_SERVICE_KEEP_ALIVE=10m \
  -e OPENMED_SERVICE_PRELOAD_MODELS=disease_detection_superclinical \
  openmed:latest

Health Check and Monitoring

The REST service exposes a /health endpoint that returns service status and version information. Use this for Docker or Kubernetes health probes:

curl http://localhost:8080/health

# {"status":"ok","service":"openmed-rest","version":"<pkg-version>","profile":"prod"}

Batch Processing for Offline Deployment

For air-gapped or offline environments where running a persistent service is unnecessary, use the BatchProcessor class to process texts directly without the HTTP layer:

from openmed import BatchProcessor, OpenMedConfig

batch = BatchProcessor(
    model_name="disease_detection_superclinical",
    group_entities=True,
    batch_size=8,
    config=OpenMedConfig(device="cpu")
)

texts = [
    "Patient has hypertension.",
    "Metastatic breast cancer treated with paclitaxel."
]

results = batch.process_texts(texts)
for r in results:
    print(r.entities)

Summary

  • Three deployment surfaces are available: Python package (openmed/__init__.py), FastAPI service (openmed/service/app.py), and Docker container (Dockerfile).
  • Model caching is handled by the ServiceRuntime in openmed/service/runtime.py, with configurable keep-alive via OPENMED_SERVICE_KEEP_ALIVE.
  • Environment variables control profile selection, preloading, and payload limits without code changes.
  • All processing is local; patient data never leaves the host when using any deployment method.

Frequently Asked Questions

How do I preload models to reduce first-request latency?

Set the OPENMED_SERVICE_PRELOAD_MODELS environment variable to a comma-separated list of model identifiers before starting the service. This triggers loading during the ServiceRuntime initialization in openmed/service/runtime.py, ensuring models are cached in memory before the first request arrives.

Can OpenMed run on Apple Silicon or iOS devices?

Yes. Install the MLX extras with pip install "openmed[mlx]" to enable the Apple Silicon backend. The openmed/core/backends.py module automatically selects the appropriate backend (PyTorch, MLX, or CoreML) based on the runtime configuration and available hardware.

Does OpenMed require an internet connection after deployment?

No. OpenMed is designed as a local-first healthcare AI library. Once models are downloaded and cached, all inference runs entirely on-device. This applies to all three deployment surfaces: Python package, REST service, and Docker container.

How do I enable GPU acceleration for OpenMed?

For CUDA devices, install the appropriate PyTorch CUDA wheels before installing OpenMed. The OpenMedConfig class accepts a device parameter (set to "cuda" or "cpu") that propagates through the ModelLoader in openmed/service/runtime.py. For Apple Silicon GPUs, use the MLX backend instead of PyTorch.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →