OpenMed REST API Configuration Options: Environment Variables, Global Settings, and Per-Request Parameters
OpenMed's REST API is configured through environment variables, a global OpenMedConfig object, and per-request Pydantic schemas that control everything from model preloading to tokenization behavior.
The maziyarpanahi/openmed repository provides a configurable inference service for medical NLP tasks. Understanding the available OpenMed REST API configuration options allows you to optimize model loading, manage memory usage, and fine-tune inference behavior for production deployments.
Service-Wide Environment Variables
OpenMed reads several environment variables at startup to configure the service runtime. These are parsed in openmed/service/runtime.py and openmed/core/config.py.
Model Preloading and Lifecycle
-
OPENMED_SERVICE_PRELOAD_MODELS– A comma-separated list of model names to load immediately when the service starts. This prevents cold-start latency for critical models. Defined inruntime.py:21-25. -
OPENMED_SERVICE_KEEP_ALIVE– The default idle timeout in seconds before a model is automatically unloaded. When a request omits thekeep_aliveparameter, this value applies. Seeruntime.py:64-70.
Configuration Profiles and Files
-
OPENMED_PROFILE– Selects a predefined configuration profile (dev,prod,test,fast, or a custom TOML profile name). This value is processed byOpenMedConfig.from_envinconfig.py:11-13. -
OPENMED_CONFIG– Path to a custom TOML configuration file, overriding the default~/.config/openmed/config.toml. Referenced inconfig.py:8-10.
Tokenizer and Authentication
-
OPENMED_USE_MEDICAL_TOKENIZER– Toggles the medical-aware tokenizer (defaults toTrue). Set to"0","false", or"no"to disable. Implemented inconfig.py:91-94. -
OPENMED_MEDICAL_TOKENIZER_EXCEPTIONS– CSV list of strings exempted from medical token remapping. Defined inconfig.py:95-98. -
HF_TOKEN– Private Hugging Face token for accessing gated models, read directly viaos.getenv. Found inconfig.py:62-64.
Global Configuration Object (OpenMedConfig)
The OpenMedConfig class in openmed/core/config.py stores service-wide settings after parsing environment variables and TOML profiles.
| Field | Description | Source Location |
|---|---|---|
default_org |
Default organization on the HuggingFace Hub | config.py:53-55 |
cache_dir |
Directory for downloaded model files (defaults to ~/.cache/openmed) |
config.py:56-58 |
device |
Preferred compute device (cpu, cuda, mps, etc.) |
config.py:59-61 |
log_level |
Logging verbosity (DEBUG, INFO, WARNING, etc.) |
config.py:65-67 |
timeout |
Global request timeout in seconds | config.py:68-70 |
use_medical_tokenizer |
Boolean flag for medical token mapper activation | config.py:71-73 |
medical_tokenizer_exceptions |
List of tokens exempt from remapping | config.py:74-76 |
backend |
Inference backend selection ("hf" for HuggingFace/PyTorch, "mlx" for Apple MLX, or None for auto-detect) |
config.py:77-79 |
profile |
Active profile name populated from OPENMED_PROFILE |
config.py:80-82 |
Runtime Configuration (ServiceRuntime)
The ServiceRuntime class manages the live service state and exposes additional configuration fields derived from environment variables.
profile– The profile string used to build the configuration.runtime.py:47-48config– TheOpenMedConfiginstance built from the profile and environment.runtime.py:48-49preload_models– Tuple of model names derived fromOPENMED_SERVICE_PRELOAD_MODELS.runtime.py:49-50default_keep_alive_seconds– Parsed value fromOPENMED_SERVICE_KEEP_ALIVE.runtime.py:50-51
Per-Request Configuration Schemas
Each API endpoint accepts a Pydantic schema defined in openmed/service/schemas.py, allowing request-level overrides of global settings.
/analyze Endpoint Options
The AnalyzeRequest schema supports the following fields (schemas.py:87-100):
text(required) – Raw clinical document textmodel_name– Model identifier (default:disease_detection_superclinical)confidence_threshold– Float between 0-1 (default:0.0)group_entities– Boolean to group adjacent entities (default:False)aggregation_strategy– One of"simple","first","average","max"(default:"simple")sentence_detection– Enable sentence splitting (default:True)sentence_language– Language code (default:"en")sentence_clean– Clean sentences before processing (default:False)use_fast_tokenizer– Use fast Rust-based tokenizer (default:True)keep_alive– Optional keep-alive specification (int, float, or string like"30s")
{
"text": "Patient presents with acute myocardial infarction...",
"model_name": "disease_detection_superclinical",
"confidence_threshold": 0.85,
"keep_alive": "5m"
}
/pii/extract Endpoint Options
The PIIExtractRequest schema (schemas.py:122-131) includes:
text(required)model_name– Defaults toOpenMed/OpenMed-PII-SuperClinical-Small-44M-v1confidence_threshold– Float 0-1 (default:0.5)use_smart_merging– Boolean (default:True)lang– Language code:en,fr,de,it,es,nl,hi,te,pt(default:en)normalize_accents– Optional booleankeep_alive– Optional keep-alive spec
/pii/deidentify Endpoint Options
The PIIDeidentifyRequest schema (schemas.py:156-170) adds deidentification-specific controls:
method–"mask","remove","replace","hash", or"shift_dates"(default:"mask")confidence_threshold– Default:0.7keep_year– Preserve year in dates (default:True)shift_dates/date_shift_days– Date shifting controls (validated together)keep_mapping– Return mapping of original to deidentified values (default:False)use_smart_merging,lang,normalize_accents,keep_alive– Same as extract
{
"text": "Patient John Doe, DOB: 1980-05-15...",
"method": "shift_dates",
"date_shift_days": 365,
"keep_mapping": true
}
/models/unload Endpoint Options
The ModelUnloadRequest schema (schemas.py:202-208) provides:
model_name– Specific model to unloadall– Boolean to unload all inactive models (default:False)
Validation ensures exactly one of these fields is provided.
Model Lifecycle and Keep-Alive Management
OpenMed uses a sophisticated idle-unload mechanism to manage GPU/CPU memory. The parse_keep_alive function in openmed/service/keep_alive.py converts human-readable strings (e.g., "30s", "5m", "2h") into float seconds values.
When a request arrives, runtime.run_model_request() (runtime.py:96-108) registers the model as active, processes the inference, and schedules an idle-unload timer via runtime._schedule_idle_unload. If no new requests arrive within the keep-alive window, runtime._unload_model_if_idle() frees the model, tokenizer, and pipeline resources.
Monitor active models via the /models/loaded endpoint (app.py:211-215), which returns cache state, active request counts, and remaining keep-alive seconds for each loaded model.
Configuration Precedence and Startup Flow
Understanding the initialization order ensures predictable behavior:
-
Environment Parsing –
create_app()inapp.py:29-47callsServiceRuntime.from_env(), which reads environment variables and constructs theOpenMedConfigbased on the selected profile. -
Model Preloading – Models listed in
OPENMED_SERVICE_PRELOAD_MODELSare loaded immediately into memory. -
Request Handling – Each endpoint validates the Pydantic schema, normalizes fields, and applies request-level
keep_alivevalues that override the service-wide default. -
Health Monitoring – The
/healthendpoint (app.py:200-209) exposes the active profile, service version, and configured timeout for operational visibility.
Summary
- Environment variables like
OPENMED_SERVICE_PRELOAD_MODELSandOPENMED_PROFILEcontrol service-wide behavior at startup - Global settings in
OpenMedConfigmanage device selection, caching, tokenization, and backend selection - Per-request schemas allow fine-tuning of confidence thresholds, aggregation strategies, and keep-alive timeouts for individual API calls
- Automatic lifecycle management unloads idle models after the specified keep-alive period to optimize memory usage
- Configuration profiles (
dev,prod,test,fast) provide quick-start presets for different deployment scenarios
Frequently Asked Questions
How do I preload specific models when starting the OpenMed service?
Set the OPENMED_SERVICE_PRELOAD_MODELS environment variable to a comma-separated list of model names before starting the server. For example: OPENMED_SERVICE_PRELOAD_MODELS=disease_detection_superclinical,OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1. This is parsed in runtime.py:21-25 and ensures models are loaded into memory at startup, eliminating cold-start latency.
Can I override the global timeout for a single API request?
While the global timeout field in OpenMedConfig sets a service-wide default, individual requests can specify a keep_alive parameter (int, float, or string like "30s" or "5m") to control how long the model remains in memory after the request completes. However, request execution timeout is controlled by the global setting. For specific timeout handling per endpoint, you must modify the global configuration or implement middleware.
What is the difference between OPENMED_PROFILE and OPENMED_CONFIG?
OPENMED_PROFILE selects a named configuration preset (dev, prod, test, fast, or custom) that maps to settings in the TOML configuration files. OPENMED_CONFIG specifies the explicit file path to a TOML configuration file, overriding the default location at ~/.config/openmed/config.toml. Use OPENMED_PROFILE to switch between predefined environments, and OPENMED_CONFIG to specify a completely custom configuration file path.
How does the automatic model unloading work?
OpenMed tracks model usage via ServiceRuntime. When a request completes, runtime._schedule_idle_unload starts a threading.Timer for the duration specified by keep_alive (request-level) or default_keep_alive_seconds (service-level). If no new requests arrive for that model before the timer expires, runtime._unload_model_if_idle removes the model from memory, freeing GPU/CPU resources. You can monitor loaded models and their remaining keep-alive time via the /models/loaded endpoint.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →