OpenMed REST API Configuration Options: Environment Variables, Global Settings, and Per-Request Parameters

OpenMed's REST API is configured through environment variables, a global OpenMedConfig object, and per-request Pydantic schemas that control everything from model preloading to tokenization behavior.

The maziyarpanahi/openmed repository provides a configurable inference service for medical NLP tasks. Understanding the available OpenMed REST API configuration options allows you to optimize model loading, manage memory usage, and fine-tune inference behavior for production deployments.

Service-Wide Environment Variables

OpenMed reads several environment variables at startup to configure the service runtime. These are parsed in openmed/service/runtime.py and openmed/core/config.py.

Model Preloading and Lifecycle

  • OPENMED_SERVICE_PRELOAD_MODELS – A comma-separated list of model names to load immediately when the service starts. This prevents cold-start latency for critical models. Defined in runtime.py:21-25.

  • OPENMED_SERVICE_KEEP_ALIVE – The default idle timeout in seconds before a model is automatically unloaded. When a request omits the keep_alive parameter, this value applies. See runtime.py:64-70.

Configuration Profiles and Files

  • OPENMED_PROFILE – Selects a predefined configuration profile (dev, prod, test, fast, or a custom TOML profile name). This value is processed by OpenMedConfig.from_env in config.py:11-13.

  • OPENMED_CONFIG – Path to a custom TOML configuration file, overriding the default ~/.config/openmed/config.toml. Referenced in config.py:8-10.

Tokenizer and Authentication

  • OPENMED_USE_MEDICAL_TOKENIZER – Toggles the medical-aware tokenizer (defaults to True). Set to "0", "false", or "no" to disable. Implemented in config.py:91-94.

  • OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS – CSV list of strings exempted from medical token remapping. Defined in config.py:95-98.

  • HF_TOKEN – Private Hugging Face token for accessing gated models, read directly via os.getenv. Found in config.py:62-64.

Global Configuration Object (OpenMedConfig)

The OpenMedConfig class in openmed/core/config.py stores service-wide settings after parsing environment variables and TOML profiles.

Field Description Source Location
default_org Default organization on the HuggingFace Hub config.py:53-55
cache_dir Directory for downloaded model files (defaults to ~/.cache/openmed) config.py:56-58
device Preferred compute device (cpu, cuda, mps, etc.) config.py:59-61
log_level Logging verbosity (DEBUG, INFO, WARNING, etc.) config.py:65-67
timeout Global request timeout in seconds config.py:68-70
use_medical_tokenizer Boolean flag for medical token mapper activation config.py:71-73
medical_tokenizer_exceptions List of tokens exempt from remapping config.py:74-76
backend Inference backend selection ("hf" for HuggingFace/PyTorch, "mlx" for Apple MLX, or None for auto-detect) config.py:77-79
profile Active profile name populated from OPENMED_PROFILE config.py:80-82

Runtime Configuration (ServiceRuntime)

The ServiceRuntime class manages the live service state and exposes additional configuration fields derived from environment variables.

  • profile – The profile string used to build the configuration. runtime.py:47-48
  • config – The OpenMedConfig instance built from the profile and environment. runtime.py:48-49
  • preload_models – Tuple of model names derived from OPENMED_SERVICE_PRELOAD_MODELS. runtime.py:49-50
  • default_keep_alive_seconds – Parsed value from OPENMED_SERVICE_KEEP_ALIVE. runtime.py:50-51

Per-Request Configuration Schemas

Each API endpoint accepts a Pydantic schema defined in openmed/service/schemas.py, allowing request-level overrides of global settings.

/analyze Endpoint Options

The AnalyzeRequest schema supports the following fields (schemas.py:87-100):

  • text (required) – Raw clinical document text
  • model_name – Model identifier (default: disease_detection_superclinical)
  • confidence_threshold – Float between 0-1 (default: 0.0)
  • group_entities – Boolean to group adjacent entities (default: False)
  • aggregation_strategy – One of "simple", "first", "average", "max" (default: "simple")
  • sentence_detection – Enable sentence splitting (default: True)
  • sentence_language – Language code (default: "en")
  • sentence_clean – Clean sentences before processing (default: False)
  • use_fast_tokenizer – Use fast Rust-based tokenizer (default: True)
  • keep_alive – Optional keep-alive specification (int, float, or string like "30s")
{
  "text": "Patient presents with acute myocardial infarction...",
  "model_name": "disease_detection_superclinical",
  "confidence_threshold": 0.85,
  "keep_alive": "5m"
}

/pii/extract Endpoint Options

The PIIExtractRequest schema (schemas.py:122-131) includes:

  • text (required)
  • model_name – Defaults to OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1
  • confidence_threshold – Float 0-1 (default: 0.5)
  • use_smart_merging – Boolean (default: True)
  • lang – Language code: en, fr, de, it, es, nl, hi, te, pt (default: en)
  • normalize_accents – Optional boolean
  • keep_alive – Optional keep-alive spec

/pii/deidentify Endpoint Options

The PIIDeidentifyRequest schema (schemas.py:156-170) adds deidentification-specific controls:

  • method – "mask", "remove", "replace", "hash", or "shift_dates" (default: "mask")
  • confidence_threshold – Default: 0.7
  • keep_year – Preserve year in dates (default: True)
  • shift_dates / date_shift_days – Date shifting controls (validated together)
  • keep_mapping – Return mapping of original to deidentified values (default: False)
  • use_smart_merging, lang, normalize_accents, keep_alive – Same as extract
{
  "text": "Patient John Doe, DOB: 1980-05-15...",
  "method": "shift_dates",
  "date_shift_days": 365,
  "keep_mapping": true
}

/models/unload Endpoint Options

The ModelUnloadRequest schema (schemas.py:202-208) provides:

  • model_name – Specific model to unload
  • all – Boolean to unload all inactive models (default: False)

Validation ensures exactly one of these fields is provided.

Model Lifecycle and Keep-Alive Management

OpenMed uses a sophisticated idle-unload mechanism to manage GPU/CPU memory. The parse_keep_alive function in openmed/service/keep_alive.py converts human-readable strings (e.g., "30s", "5m", "2h") into float seconds values.

When a request arrives, runtime.run_model_request() (runtime.py:96-108) registers the model as active, processes the inference, and schedules an idle-unload timer via runtime._schedule_idle_unload. If no new requests arrive within the keep-alive window, runtime._unload_model_if_idle() frees the model, tokenizer, and pipeline resources.

Monitor active models via the /models/loaded endpoint (app.py:211-215), which returns cache state, active request counts, and remaining keep-alive seconds for each loaded model.

Configuration Precedence and Startup Flow

Understanding the initialization order ensures predictable behavior:

  1. Environment Parsing – create_app() in app.py:29-47 calls ServiceRuntime.from_env(), which reads environment variables and constructs the OpenMedConfig based on the selected profile.

  2. Model Preloading – Models listed in OPENMED_SERVICE_PRELOAD_MODELS are loaded immediately into memory.

  3. Request Handling – Each endpoint validates the Pydantic schema, normalizes fields, and applies request-level keep_alive values that override the service-wide default.

  4. Health Monitoring – The /health endpoint (app.py:200-209) exposes the active profile, service version, and configured timeout for operational visibility.

Summary

  • Environment variables like OPENMED_SERVICE_PRELOAD_MODELS and OPENMED_PROFILE control service-wide behavior at startup
  • Global settings in OpenMedConfig manage device selection, caching, tokenization, and backend selection
  • Per-request schemas allow fine-tuning of confidence thresholds, aggregation strategies, and keep-alive timeouts for individual API calls
  • Automatic lifecycle management unloads idle models after the specified keep-alive period to optimize memory usage
  • Configuration profiles (dev, prod, test, fast) provide quick-start presets for different deployment scenarios

Frequently Asked Questions

How do I preload specific models when starting the OpenMed service?

Set the OPENMED_SERVICE_PRELOAD_MODELS environment variable to a comma-separated list of model names before starting the server. For example: OPENMED_SERVICE_PRELOAD_MODELS=disease_detection_superclinical,OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1. This is parsed in runtime.py:21-25 and ensures models are loaded into memory at startup, eliminating cold-start latency.

Can I override the global timeout for a single API request?

While the global timeout field in OpenMedConfig sets a service-wide default, individual requests can specify a keep_alive parameter (int, float, or string like "30s" or "5m") to control how long the model remains in memory after the request completes. However, request execution timeout is controlled by the global setting. For specific timeout handling per endpoint, you must modify the global configuration or implement middleware.

What is the difference between OPENMED_PROFILE and OPENMED_CONFIG?

OPENMED_PROFILE selects a named configuration preset (dev, prod, test, fast, or custom) that maps to settings in the TOML configuration files. OPENMED_CONFIG specifies the explicit file path to a TOML configuration file, overriding the default location at ~/.config/openmed/config.toml. Use OPENMED_PROFILE to switch between predefined environments, and OPENMED_CONFIG to specify a completely custom configuration file path.

How does the automatic model unloading work?

OpenMed tracks model usage via ServiceRuntime. When a request completes, runtime._schedule_idle_unload starts a threading.Timer for the duration specified by keep_alive (request-level) or default_keep_alive_seconds (service-level). If no new requests arrive for that model before the timer expires, runtime._unload_model_if_idle removes the model from memory, freeing GPU/CPU resources. You can monitor loaded models and their remaining keep-alive time via the /models/loaded endpoint.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →