# How to Use OpenMed's REST Service with FastAPI: A Complete Guide

> Learn to use OpenMed's REST service with FastAPI. This guide covers clinical NLP, PII extraction, and model management via a simple HTTP interface. Integrate powerful text analysis into your applications today.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**OpenMed provides a production-ready REST API built on FastAPI that exposes clinical NLP endpoints for text analysis, PII extraction, and model management through a simple HTTP interface.**

OpenMed ships with a fully-featured REST service built on FastAPI, enabling seamless integration of clinical NLP pipelines into existing healthcare applications. The service architecture in [`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py) provides a robust runtime environment with automatic model loading, configurable timeouts, and standardized error handling. This guide covers how to start the server, interact with available endpoints, and handle responses according to the source code implementation in the `maziyarpanahi/openmed` repository.

## Architecture Overview

The FastAPI application follows a layered architecture centered around a `ServiceRuntime` object that manages model lifecycle and configuration.

**Application bootstrap** occurs in the `create_app()` function within [`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py). This factory function instantiates a `FastAPI` object, registers a custom *lifespan* hook that initializes the `ServiceRuntime` from environment variables, and pre-loads models when configured. The runtime is stored on the FastAPI `app.state` object via `_attach_runtime` and accessed lazily through `_get_service_runtime` on first request.

**Timeout management** ensures reliability for heavy model inference. All blocking model calls execute in a thread-pool and are wrapped by `_run_with_timeout`, which respects the `timeout` value from the runtime configuration and raises a `ServiceTimeoutError` when limits are exceeded.

**Error handling** translates FastAPI, Starlette, and internal validation errors into uniform JSON responses using the `_error_response` helper, ensuring consistent API contracts across failure modes.

## Starting the Server

Launch the REST service using **uvicorn** with the auto-generated `app` object at the bottom of the service module:

```bash
uvicorn openmed.service.app:app --host 0.0.0.0 --port 8000

```

Replace `--host` and `--port` values as needed for your deployment environment. The server initializes the `ServiceRuntime` during startup, loading any models specified in your environment configuration before accepting traffic.

## Core Endpoints and Usage

The API exposes endpoints for health monitoring, clinical text analysis, PII processing, and model management. Each endpoint extracts the runtime from `app.state`, constructs a blocking operation calling OpenMed’s high-level functions, and executes it via `_run_with_timeout`.

### Health Monitoring

Verify service status and configuration:

```bash
curl http://localhost:8000/health

```

The response returns service metadata:

```json
{
  "status": "ok",
  "service": "openmed-rest",
  "version": "X.Y.Z",
  "profile": "default"
}

```

### Text Analysis

Submit clinical text for entity recognition and analysis using the `POST /analyze` endpoint. The endpoint accepts an `AnalyzeRequest` payload as defined in [`openmed/service/schemas.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/schemas.py):

```python
import requests

payload = {
    "text": "Patient was diagnosed with diabetes mellitus.",
    "model_name": "clinical_ner",
    "keep_alive": True,
    "aggregation_strategy": "average",
    "confidence_threshold": 0.5,
    "group_entities": False,
    "sentence_detection": True,
    "sentence_language": "en",
    "sentence_clean": True,
    "use_fast_tokenizer": True,
}

resp = requests.post("http://localhost:8000/analyze", json=payload)
print(resp.json())

```

### PII Processing

Extract personally identifiable information using `POST /pii/extract`:

```bash
curl -X POST http://localhost:8000/pii/extract \
     -H "Content-Type: application/json" \
     -d '{
           "text": "John Doe, SSN 123-45-6789, visited on 2021-03-15.",
           "model_name": "pii_en",
           "confidence_threshold": 0.6,
           "use_smart_merging": true,
           "lang": "en",
           "normalize_accents": false
         }'

```

De-identify text using `POST /pii/deidentify` with configurable masking strategies:

```python
payload = {
    "text": "John Doe, SSN 123-45-6789, visited on 2021-03-15.",
    "method": "mask",
    "model_name": "pii_en",
    "confidence_threshold": 0.6,
    "keep_year": False,
    "shift_dates": False,
    "date_shift_days": 0,
    "keep_mapping": False,
    "use_smart_merging": True,
    "lang": "en",
    "normalize_accents": False,
}
resp = requests.post("http://localhost:8000/pii/deidentify", json=payload)
print(resp.json())

```

### Model Management

Inspect currently loaded models to monitor memory usage:

```bash
curl http://localhost:8000/models/loaded

```

Unload specific models to free resources without restarting the service:

```bash
curl -X POST http://localhost:8000/models/unload \
     -H "Content-Type: application/json" \
     -d '{"model_name": "clinical_ner", "all": false}'

```

Set `"all": true` to unload all models simultaneously.

## Error Handling and Timeouts

The REST service implements comprehensive error handling through FastAPI exception handlers registered in [`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py). All endpoints wrap model inference calls in `_run_with_timeout`, which executes heavy operations in a thread-pool and enforces the configured timeout limit.

When timeouts occur, the API returns a structured error response generated by `_error_response`, maintaining the JSON envelope format used for successful responses. Validation errors from Pydantic models in [`openmed/service/schemas.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/schemas.py) return detailed field-level error messages, while internal failures return sanitized error codes to prevent information leakage.

## Summary

- **OpenMed's REST service** is built on FastAPI and defined in [`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py), providing a production-ready HTTP interface for clinical NLP.
- **Start the server** using `uvicorn openmed.service.app:app` with configurable host and port bindings.
- **Core endpoints** include `/health` for monitoring, `/analyze` for clinical text processing, `/pii/extract` and `/pii/deidentify` for privacy operations, and `/models/loaded` and `/models/unload` for runtime management.
- **Timeout protection** via `_run_with_timeout` prevents hanging requests during model inference by respecting runtime configuration limits.
- **Consistent error responses** use `_error_response` to standardize JSON output across validation failures, timeouts, and internal errors.

## Frequently Asked Questions

### How do I start the OpenMed REST service locally?

Execute `uvicorn openmed.service.app:app --host 0.0.0.0 --port 8000` from your Python environment. The `create_app()` factory function in [`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py) initializes the FastAPI instance and `ServiceRuntime` automatically, loading any pre-configured models before accepting connections.

### What is the ServiceRuntime and how does it work?

The `ServiceRuntime` is a stateful object created from environment variables that manages model loading, configuration, and resource allocation. It attaches to the FastAPI `app.state` via `_attach_runtime` and is accessed lazily through `_get_service_runtime`, ensuring thread-safe initialization on the first request while preserving runtime context across subsequent calls.

### How does the API handle long-running model requests?

All heavy model calls execute through `_run_with_timeout`, which runs inference in a thread-pool and enforces the `timeout` value specified in the runtime configuration. If execution exceeds this limit, the function raises a `ServiceTimeoutError`, which the global exception handler converts into a standardized JSON error response.

### Which endpoints are available for PII processing?

The API exposes two primary PII endpoints in [`openmed/service/app.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/app.py): `POST /pii/extract` for identifying sensitive entities like names, SSNs, and dates, and `POST /pii/deidentify` for masking or anonymizing that information. Both endpoints accept model selection, confidence thresholds, and language parameters as defined in the Pydantic schemas.