# How to Deploy OpenMed: Complete Guide to Python Package, REST API, and Docker Installation

> Deploy OpenMed with pip, FastAPI, or Docker. Install the Python package, run the REST API, or build the Docker image to serve local-first healthcare AI.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-13

---

**Deploy OpenMed by installing the Python package with `pip install "openmed[hf]"`, running the FastAPI REST service with `uvicorn openmed.service.app:app`, or building the Docker image from the repository's Dockerfile to serve the local-first healthcare AI on any infrastructure.**

OpenMed is a local-first healthcare AI library developed by maziyarpanahi/openmed that runs completely on-device, ensuring patient data never leaves the host. The repository provides three deployment surfaces: a **Python package** for direct library integration, a **FastAPI REST service** defined in [`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py), and a **production-ready Docker container**. This guide covers each deployment method using the actual source code structure and configuration options found in the repository.

## Install OpenMed as a Python Package

The quickest way to deploy OpenMed is via `pip` or `uv`. The library exposes its public API through [`openmed/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/__init__.py), which includes functions like `analyze_text`, `extract_pii`, and `deidentify`.

Install the core library with Hugging Face extras for CPU and CUDA support:

```bash
pip install "openmed[hf]"

```

For Apple Silicon acceleration using the MLX backend, reference the backend selection logic in [`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py) and install the MLX variant:

```bash
pip install "openmed[mlx]"

```

After installation, you can call the high-level API directly:

```python
from openmed import analyze_text

result = analyze_text(
    "Patient started on imatinib for chronic myeloid leukemia.",
    model_name="disease_detection_superclinical"
)

print(result.entities)

```

## Run the OpenMed REST Service Locally

The REST service is a FastAPI application defined in [`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py) that exposes HTTP endpoints for `/analyze`, `/pii/extract`, and `/pii/deidentify`. The service uses a shared `ServiceRuntime` instantiated in [`openmed/service/runtime.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/runtime.py) to manage model loading and keep-alive policies across requests.

### Install Service Dependencies

Install the optional service extra to include FastAPI, uvicorn, and Pydantic models from [`openmed/service/schemas.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/schemas.py):

```bash
pip install "openmed[hf,service]"

```

### Start the Server

Launch the service using uvicorn with the app factory defined in [`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py):

```bash
uvicorn openmed.service.app:app --host 0.0.0.0 --port 8080

```

### Configure Environment Variables

The runtime behavior is controlled through environment variables read by the `ServiceRuntime` class. Set these before starting the server:

- **OPENMED_PROFILE**: Selects runtime profile (`dev`, `prod`, `test`) affecting request timeouts
- **OPENMED_SERVICE_KEEP_ALIVE**: Model cache TTL (e.g., `10m`); set to `0` to disable caching
- **OPENMED_SERVICE_PRELOAD_MODELS**: Comma-separated list of models to load at startup (e.g., `disease_detection_superclinical,OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1`)
- **OPENMED_SERVICE_MAX_TEXT_LENGTH**: Maximum request payload size for pre-validation (default `250000`)

Example request using curl:

```bash
curl -X POST http://127.0.0.1:8080/analyze \
  -H "Content-Type: application/json" \
  -d '{"text":"Patient started imatinib for CML.", "model_name":"disease_detection_superclinical"}'

```

## Deploy OpenMed with Docker

For production deployments, the repository includes a `Dockerfile` that builds a minimal Python-slim image with the service pre-installed.

### Build the Image

```bash
docker build -t openmed:latest .

```

### Run the Container

The container automatically starts `uvicorn openmed.service.app:app` via the `CMD` instruction in the Dockerfile. Pass environment variables to configure the runtime:

```bash
docker run --rm -p 8080:8080 \
  -e OPENMED_PROFILE=prod \
  -e OPENMED_SERVICE_KEEP_ALIVE=10m \
  -e OPENMED_SERVICE_PRELOAD_MODELS=disease_detection_superclinical \
  openmed:latest

```

### Health Check and Monitoring

The REST service exposes a `/health` endpoint that returns service status and version information. Use this for Docker or Kubernetes health probes:

```bash
curl http://localhost:8080/health

# {"status":"ok","service":"openmed-rest","version":"<pkg-version>","profile":"prod"}

```

## Batch Processing for Offline Deployment

For air-gapped or offline environments where running a persistent service is unnecessary, use the `BatchProcessor` class to process texts directly without the HTTP layer:

```python
from openmed import BatchProcessor, OpenMedConfig

batch = BatchProcessor(
    model_name="disease_detection_superclinical",
    group_entities=True,
    batch_size=8,
    config=OpenMedConfig(device="cpu")
)

texts = [
    "Patient has hypertension.",
    "Metastatic breast cancer treated with paclitaxel."
]

results = batch.process_texts(texts)
for r in results:
    print(r.entities)

```

## Summary

- **Three deployment surfaces** are available: Python package ([`openmed/__init__.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/__init__.py)), FastAPI service ([`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py)), and Docker container (`Dockerfile`).
- **Model caching** is handled by the `ServiceRuntime` in [`openmed/service/runtime.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/runtime.py), with configurable keep-alive via `OPENMED_SERVICE_KEEP_ALIVE`.
- **Environment variables** control profile selection, preloading, and payload limits without code changes.
- **All processing is local**; patient data never leaves the host when using any deployment method.

## Frequently Asked Questions

### How do I preload models to reduce first-request latency?

Set the `OPENMED_SERVICE_PRELOAD_MODELS` environment variable to a comma-separated list of model identifiers before starting the service. This triggers loading during the `ServiceRuntime` initialization in [`openmed/service/runtime.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/runtime.py), ensuring models are cached in memory before the first request arrives.

### Can OpenMed run on Apple Silicon or iOS devices?

Yes. Install the MLX extras with `pip install "openmed[mlx]"` to enable the Apple Silicon backend. The [`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py) module automatically selects the appropriate backend (PyTorch, MLX, or CoreML) based on the runtime configuration and available hardware.

### Does OpenMed require an internet connection after deployment?

No. OpenMed is designed as a local-first healthcare AI library. Once models are downloaded and cached, all inference runs entirely on-device. This applies to all three deployment surfaces: Python package, REST service, and Docker container.

### How do I enable GPU acceleration for OpenMed?

For CUDA devices, install the appropriate PyTorch CUDA wheels before installing OpenMed. The `OpenMedConfig` class accepts a `device` parameter (set to `"cuda"` or `"cpu"`) that propagates through the `ModelLoader` in [`openmed/service/runtime.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/runtime.py). For Apple Silicon GPUs, use the MLX backend instead of PyTorch.