# OpenMed REST API Configuration Options: Environment Variables, Global Settings, and Per-Request Parameters

> Discover OpenMed REST API configuration: master environment variables, global settings, and per-request parameters for fine-grained control over model preloading and tokenization.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: api-reference
- Published: 2026-06-10

---

**OpenMed's REST API is configured through environment variables, a global `OpenMedConfig` object, and per-request Pydantic schemas that control everything from model preloading to tokenization behavior.**

The `maziyarpanahi/openmed` repository provides a configurable inference service for medical NLP tasks. Understanding the available **OpenMed REST API configuration options** allows you to optimize model loading, manage memory usage, and fine-tune inference behavior for production deployments.

## Service-Wide Environment Variables

OpenMed reads several environment variables at startup to configure the service runtime. These are parsed in [`openmed/service/runtime.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/runtime.py) and [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py).

### Model Preloading and Lifecycle

- **`OPENMED_SERVICE_PRELOAD_MODELS`** – A comma-separated list of model names to load immediately when the service starts. This prevents cold-start latency for critical models. Defined in [`runtime.py:21-25`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/runtime.py#L21-L25).

- **`OPENMED_SERVICE_KEEP_ALIVE`** – The default idle timeout in seconds before a model is automatically unloaded. When a request omits the `keep_alive` parameter, this value applies. See [`runtime.py:64-70`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/runtime.py#L64-L70).

### Configuration Profiles and Files

- **`OPENMED_PROFILE`** – Selects a predefined configuration profile (`dev`, `prod`, `test`, `fast`, or a custom TOML profile name). This value is processed by `OpenMedConfig.from_env` in [`config.py:11-13`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L11-L13).

- **`OPENMED_CONFIG`** – Path to a custom TOML configuration file, overriding the default `~/.config/openmed/config.toml`. Referenced in [`config.py:8-10`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L8-L10).

### Tokenizer and Authentication

- **`OPENMED_USE_MEDICAL_TOKENIZER`** – Toggles the medical-aware tokenizer (defaults to `True`). Set to `"0"`, `"false"`, or `"no"` to disable. Implemented in [`config.py:91-94`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L91-L94).

- **`OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS`** – CSV list of strings exempted from medical token remapping. Defined in [`config.py:95-98`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L95-L98).

- **`HF_TOKEN`** – Private Hugging Face token for accessing gated models, read directly via `os.getenv`. Found in [`config.py:62-64`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L62-L64).

## Global Configuration Object (OpenMedConfig)

The `OpenMedConfig` class in [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py) stores service-wide settings after parsing environment variables and TOML profiles.

| Field | Description | Source Location |
|-------|-------------|-----------------|
| `default_org` | Default organization on the HuggingFace Hub | [`config.py:53-55`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L53-L55) |
| `cache_dir` | Directory for downloaded model files (defaults to `~/.cache/openmed`) | [`config.py:56-58`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L56-L58) |
| `device` | Preferred compute device (`cpu`, `cuda`, `mps`, etc.) | [`config.py:59-61`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L59-L61) |
| `log_level` | Logging verbosity (`DEBUG`, `INFO`, `WARNING`, etc.) | [`config.py:65-67`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L65-L67) |
| `timeout` | Global request timeout in seconds | [`config.py:68-70`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L68-L70) |
| `use_medical_tokenizer` | Boolean flag for medical token mapper activation | [`config.py:71-73`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L71-L73) |
| `medical_tokenizer_exceptions` | List of tokens exempt from remapping | [`config.py:74-76`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L74-L76) |
| `backend` | Inference backend selection (`"hf"` for HuggingFace/PyTorch, `"mlx"` for Apple MLX, or `None` for auto-detect) | [`config.py:77-79`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L77-L79) |
| `profile` | Active profile name populated from `OPENMED_PROFILE` | [`config.py:80-82`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/core/config.py#L80-L82) |

## Runtime Configuration (ServiceRuntime)

The `ServiceRuntime` class manages the live service state and exposes additional configuration fields derived from environment variables.

- **`profile`** – The profile string used to build the configuration. [`runtime.py:47-48`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/runtime.py#L47-L48)
- **`config`** – The `OpenMedConfig` instance built from the profile and environment. [`runtime.py:48-49`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/runtime.py#L48-L49)
- **`preload_models`** – Tuple of model names derived from `OPENMED_SERVICE_PRELOAD_MODELS`. [`runtime.py:49-50`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/runtime.py#L49-L50)
- **`default_keep_alive_seconds`** – Parsed value from `OPENMED_SERVICE_KEEP_ALIVE`. [`runtime.py:50-51`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/runtime.py#L50-L51)

## Per-Request Configuration Schemas

Each API endpoint accepts a Pydantic schema defined in [`openmed/service/schemas.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/schemas.py), allowing request-level overrides of global settings.

### /analyze Endpoint Options

The `AnalyzeRequest` schema supports the following fields ([`schemas.py:87-100`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/schemas.py#L87-L100)):

- **`text`** (required) – Raw clinical document text
- **`model_name`** – Model identifier (default: `disease_detection_superclinical`)
- **`confidence_threshold`** – Float between 0-1 (default: `0.0`)
- **`group_entities`** – Boolean to group adjacent entities (default: `False`)
- **`aggregation_strategy`** – One of `"simple"`, `"first"`, `"average"`, `"max"` (default: `"simple"`)
- **`sentence_detection`** – Enable sentence splitting (default: `True`)
- **`sentence_language`** – Language code (default: `"en"`)
- **`sentence_clean`** – Clean sentences before processing (default: `False`)
- **`use_fast_tokenizer`** – Use fast Rust-based tokenizer (default: `True`)
- **`keep_alive`** – Optional keep-alive specification (int, float, or string like `"30s"`)

```json
{
  "text": "Patient presents with acute myocardial infarction...",
  "model_name": "disease_detection_superclinical",
  "confidence_threshold": 0.85,
  "keep_alive": "5m"
}

```

### /pii/extract Endpoint Options

The `PIIExtractRequest` schema ([`schemas.py:122-131`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/schemas.py#L122-L131)) includes:

- **`text`** (required)
- **`model_name`** – Defaults to `OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1`
- **`confidence_threshold`** – Float 0-1 (default: `0.5`)
- **`use_smart_merging`** – Boolean (default: `True`)
- **`lang`** – Language code: `en`, `fr`, `de`, `it`, `es`, `nl`, `hi`, `te`, `pt` (default: `en`)
- **`normalize_accents`** – Optional boolean
- **`keep_alive`** – Optional keep-alive spec

### /pii/deidentify Endpoint Options

The `PIIDeidentifyRequest` schema ([`schemas.py:156-170`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/schemas.py#L156-L170)) adds deidentification-specific controls:

- **`method`** – `"mask"`, `"remove"`, `"replace"`, `"hash"`, or `"shift_dates"` (default: `"mask"`)
- **`confidence_threshold`** – Default: `0.7`
- **`keep_year`** – Preserve year in dates (default: `True`)
- **`shift_dates`** / **`date_shift_days`** – Date shifting controls (validated together)
- **`keep_mapping`** – Return mapping of original to deidentified values (default: `False`)
- **`use_smart_merging`**, **`lang`**, **`normalize_accents`**, **`keep_alive`** – Same as extract

```json
{
  "text": "Patient John Doe, DOB: 1980-05-15...",
  "method": "shift_dates",
  "date_shift_days": 365,
  "keep_mapping": true
}

```

### /models/unload Endpoint Options

The `ModelUnloadRequest` schema ([`schemas.py:202-208`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/schemas.py#L202-L208)) provides:

- **`model_name`** – Specific model to unload
- **`all`** – Boolean to unload all inactive models (default: `False`)

Validation ensures exactly one of these fields is provided.

## Model Lifecycle and Keep-Alive Management

OpenMed uses a sophisticated idle-unload mechanism to manage GPU/CPU memory. The `parse_keep_alive` function in [`openmed/service/keep_alive.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/service/keep_alive.py) converts human-readable strings (e.g., `"30s"`, `"5m"`, `"2h"`) into float seconds values.

When a request arrives, `runtime.run_model_request()` ([`runtime.py:96-108`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/runtime.py#L96-L108)) registers the model as active, processes the inference, and schedules an idle-unload timer via `runtime._schedule_idle_unload`. If no new requests arrive within the keep-alive window, `runtime._unload_model_if_idle()` frees the model, tokenizer, and pipeline resources.

Monitor active models via the `/models/loaded` endpoint ([`app.py:211-215`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/app.py#L211-L215)), which returns cache state, active request counts, and remaining keep-alive seconds for each loaded model.

## Configuration Precedence and Startup Flow

Understanding the initialization order ensures predictable behavior:

1. **Environment Parsing** – `create_app()` in [`app.py:29-47`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/app.py#L29-L47) calls `ServiceRuntime.from_env()`, which reads environment variables and constructs the `OpenMedConfig` based on the selected profile.

2. **Model Preloading** – Models listed in `OPENMED_SERVICE_PRELOAD_MODELS` are loaded immediately into memory.

3. **Request Handling** – Each endpoint validates the Pydantic schema, normalizes fields, and applies request-level `keep_alive` values that override the service-wide default.

4. **Health Monitoring** – The `/health` endpoint ([`app.py:200-209`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/app.py#L200-L209)) exposes the active profile, service version, and configured timeout for operational visibility.

## Summary

- **Environment variables** like `OPENMED_SERVICE_PRELOAD_MODELS` and `OPENMED_PROFILE` control service-wide behavior at startup
- **Global settings** in `OpenMedConfig` manage device selection, caching, tokenization, and backend selection
- **Per-request schemas** allow fine-tuning of confidence thresholds, aggregation strategies, and keep-alive timeouts for individual API calls
- **Automatic lifecycle management** unloads idle models after the specified keep-alive period to optimize memory usage
- **Configuration profiles** (`dev`, `prod`, `test`, `fast`) provide quick-start presets for different deployment scenarios

## Frequently Asked Questions

### How do I preload specific models when starting the OpenMed service?

Set the `OPENMED_SERVICE_PRELOAD_MODELS` environment variable to a comma-separated list of model names before starting the server. For example: `OPENMED_SERVICE_PRELOAD_MODELS=disease_detection_superclinical,OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1`. This is parsed in [`runtime.py:21-25`](https://github.com/maziyarpanahi/openmed/blob/master/openmed/service/runtime.py#L21-L25) and ensures models are loaded into memory at startup, eliminating cold-start latency.

### Can I override the global timeout for a single API request?

While the global `timeout` field in `OpenMedConfig` sets a service-wide default, individual requests can specify a `keep_alive` parameter (int, float, or string like `"30s"` or `"5m"`) to control how long the model remains in memory after the request completes. However, request execution timeout is controlled by the global setting. For specific timeout handling per endpoint, you must modify the global configuration or implement middleware.

### What is the difference between OPENMED_PROFILE and OPENMED_CONFIG?

`OPENMED_PROFILE` selects a named configuration preset (`dev`, `prod`, `test`, `fast`, or custom) that maps to settings in the TOML configuration files. `OPENMED_CONFIG` specifies the explicit file path to a TOML configuration file, overriding the default location at `~/.config/openmed/config.toml`. Use `OPENMED_PROFILE` to switch between predefined environments, and `OPENMED_CONFIG` to specify a completely custom configuration file path.

### How does the automatic model unloading work?

OpenMed tracks model usage via `ServiceRuntime`. When a request completes, `runtime._schedule_idle_unload` starts a `threading.Timer` for the duration specified by `keep_alive` (request-level) or `default_keep_alive_seconds` (service-level). If no new requests arrive for that model before the timer expires, `runtime._unload_model_if_idle` removes the model from memory, freeing GPU/CPU resources. You can monitor loaded models and their remaining keep-alive time via the `/models/loaded` endpoint.