How to Use OpenMed's REST Service with FastAPI: A Complete Guide
OpenMed provides a production-ready REST API built on FastAPI that exposes clinical NLP endpoints for text analysis, PII extraction, and model management through a simple HTTP interface.
OpenMed ships with a fully-featured REST service built on FastAPI, enabling seamless integration of clinical NLP pipelines into existing healthcare applications. The service architecture in openmed/service/app.py provides a robust runtime environment with automatic model loading, configurable timeouts, and standardized error handling. This guide covers how to start the server, interact with available endpoints, and handle responses according to the source code implementation in the maziyarpanahi/openmed repository.
Architecture Overview
The FastAPI application follows a layered architecture centered around a ServiceRuntime object that manages model lifecycle and configuration.
Application bootstrap occurs in the create_app() function within openmed/service/app.py. This factory function instantiates a FastAPI object, registers a custom lifespan hook that initializes the ServiceRuntime from environment variables, and pre-loads models when configured. The runtime is stored on the FastAPI app.state object via _attach_runtime and accessed lazily through _get_service_runtime on first request.
Timeout management ensures reliability for heavy model inference. All blocking model calls execute in a thread-pool and are wrapped by _run_with_timeout, which respects the timeout value from the runtime configuration and raises a ServiceTimeoutError when limits are exceeded.
Error handling translates FastAPI, Starlette, and internal validation errors into uniform JSON responses using the _error_response helper, ensuring consistent API contracts across failure modes.
Starting the Server
Launch the REST service using uvicorn with the auto-generated app object at the bottom of the service module:
uvicorn openmed.service.app:app --host 0.0.0.0 --port 8000
Replace --host and --port values as needed for your deployment environment. The server initializes the ServiceRuntime during startup, loading any models specified in your environment configuration before accepting traffic.
Core Endpoints and Usage
The API exposes endpoints for health monitoring, clinical text analysis, PII processing, and model management. Each endpoint extracts the runtime from app.state, constructs a blocking operation calling OpenMed’s high-level functions, and executes it via _run_with_timeout.
Health Monitoring
Verify service status and configuration:
curl http://localhost:8000/health
The response returns service metadata:
{
"status": "ok",
"service": "openmed-rest",
"version": "X.Y.Z",
"profile": "default"
}
Text Analysis
Submit clinical text for entity recognition and analysis using the POST /analyze endpoint. The endpoint accepts an AnalyzeRequest payload as defined in openmed/service/schemas.py:
import requests
payload = {
"text": "Patient was diagnosed with diabetes mellitus.",
"model_name": "clinical_ner",
"keep_alive": True,
"aggregation_strategy": "average",
"confidence_threshold": 0.5,
"group_entities": False,
"sentence_detection": True,
"sentence_language": "en",
"sentence_clean": True,
"use_fast_tokenizer": True,
}
resp = requests.post("http://localhost:8000/analyze", json=payload)
print(resp.json())
PII Processing
Extract personally identifiable information using POST /pii/extract:
curl -X POST http://localhost:8000/pii/extract \
-H "Content-Type: application/json" \
-d '{
"text": "John Doe, SSN 123-45-6789, visited on 2021-03-15.",
"model_name": "pii_en",
"confidence_threshold": 0.6,
"use_smart_merging": true,
"lang": "en",
"normalize_accents": false
}'
De-identify text using POST /pii/deidentify with configurable masking strategies:
payload = {
"text": "John Doe, SSN 123-45-6789, visited on 2021-03-15.",
"method": "mask",
"model_name": "pii_en",
"confidence_threshold": 0.6,
"keep_year": False,
"shift_dates": False,
"date_shift_days": 0,
"keep_mapping": False,
"use_smart_merging": True,
"lang": "en",
"normalize_accents": False,
}
resp = requests.post("http://localhost:8000/pii/deidentify", json=payload)
print(resp.json())
Model Management
Inspect currently loaded models to monitor memory usage:
curl http://localhost:8000/models/loaded
Unload specific models to free resources without restarting the service:
curl -X POST http://localhost:8000/models/unload \
-H "Content-Type: application/json" \
-d '{"model_name": "clinical_ner", "all": false}'
Set "all": true to unload all models simultaneously.
Error Handling and Timeouts
The REST service implements comprehensive error handling through FastAPI exception handlers registered in openmed/service/app.py. All endpoints wrap model inference calls in _run_with_timeout, which executes heavy operations in a thread-pool and enforces the configured timeout limit.
When timeouts occur, the API returns a structured error response generated by _error_response, maintaining the JSON envelope format used for successful responses. Validation errors from Pydantic models in openmed/service/schemas.py return detailed field-level error messages, while internal failures return sanitized error codes to prevent information leakage.
Summary
- OpenMed's REST service is built on FastAPI and defined in
openmed/service/app.py, providing a production-ready HTTP interface for clinical NLP. - Start the server using
uvicorn openmed.service.app:appwith configurable host and port bindings. - Core endpoints include
/healthfor monitoring,/analyzefor clinical text processing,/pii/extractand/pii/deidentifyfor privacy operations, and/models/loadedand/models/unloadfor runtime management. - Timeout protection via
_run_with_timeoutprevents hanging requests during model inference by respecting runtime configuration limits. - Consistent error responses use
_error_responseto standardize JSON output across validation failures, timeouts, and internal errors.
Frequently Asked Questions
How do I start the OpenMed REST service locally?
Execute uvicorn openmed.service.app:app --host 0.0.0.0 --port 8000 from your Python environment. The create_app() factory function in openmed/service/app.py initializes the FastAPI instance and ServiceRuntime automatically, loading any pre-configured models before accepting connections.
What is the ServiceRuntime and how does it work?
The ServiceRuntime is a stateful object created from environment variables that manages model loading, configuration, and resource allocation. It attaches to the FastAPI app.state via _attach_runtime and is accessed lazily through _get_service_runtime, ensuring thread-safe initialization on the first request while preserving runtime context across subsequent calls.
How does the API handle long-running model requests?
All heavy model calls execute through _run_with_timeout, which runs inference in a thread-pool and enforces the timeout value specified in the runtime configuration. If execution exceeds this limit, the function raises a ServiceTimeoutError, which the global exception handler converts into a standardized JSON error response.
Which endpoints are available for PII processing?
The API exposes two primary PII endpoints in openmed/service/app.py: POST /pii/extract for identifying sensitive entities like names, SSNs, and dates, and POST /pii/deidentify for masking or anonymizing that information. Both endpoints accept model selection, confidence thresholds, and language parameters as defined in the Pydantic schemas.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →