Database Schema for OpenMed: Understanding Pydantic Request Models

The OpenMed database schema consists of four Pydantic request models—AnalyzeRequest, PIIExtractRequest, PIIDeidentifyRequest, and ModelUnloadRequest—that define strict validation rules and field structures for the REST-MCP service, replacing traditional relational database tables with type-safe Python contracts.

OpenMed (maziyarpanahi/openmed) does not expose a traditional relational database schema. Instead, the core data architecture relies on Pydantic request schemas defined in openmed/service/schemas.py that enforce validation, type safety, and field defaults for all client-server interactions. These schemas act as the authoritative contract between clients and the OpenMed server, ensuring every payload meets strict structural requirements before processing.

Core Request Schema Architecture

The database schema for OpenMed centers on four primary Pydantic models that govern how data enters the system. Each model inherits from _StrictModel, a base class configured with extra="forbid" to reject unknown fields immediately. This design eliminates silent failures by enforcing explicit contracts at the API boundary.

The four primary request models are:

  • AnalyzeRequest – Handles clinical text analysis and entity extraction
  • PIIExtractRequest – Manages personally identifiable information extraction
  • PIIDeidentifyRequest – Controls PII masking, replacement, and date shifting
  • ModelUnloadRequest – Governs model lifecycle management

Primary Request Models

AnalyzeRequest

The AnalyzeRequest schema in openmed/service/schemas.py (lines 103-115) defines the structure for general text analysis operations including clinical entity extraction and classification.

Key fields:

  • text: str – The raw clinical document, normalized via _normalize_text and length-checked against get_max_text_length
  • model_name: str – Defaults to "disease_detection_superclinical", validated through validate_model_name
  • confidence_threshold: float – Clamped to 0-1 range via validate_confidence_threshold, defaults to 0.0
  • group_entities: bool – Defaults to False
  • aggregation_strategy: Literal["simple","first","average","max"] – Defaults to "simple"
  • sentence_detection: bool – Defaults to True
  • sentence_language: str – Defaults to "en"
  • sentence_clean: bool – Defaults to False
  • use_fast_tokenizer: bool – Defaults to True
  • keep_alive: KeepAliveValue – Optional; accepts seconds, "inf", or "auto"

PIIExtractRequest

The PIIExtractRequest schema (lines 139-148) structures payloads for extracting personally identifiable information from clinical documents.

Key fields:

  • text: str – Normalized using the same _normalize_text pipeline as AnalyzeRequest
  • model_name: str – Defaults to _DEFAULT_PII_MODEL ("OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1")
  • confidence_threshold: float – Defaults to 0.5
  • use_smart_merging: bool – Defaults to True
  • lang: PIILanguage – Enum literal defaulting to "en"; must be one of: "en", "fr", "de", "it", "es", "nl", "hi", "te", "pt", "ar", "ja", "tr"
  • normalize_accents: Optional[bool] – Optional accent normalization
  • keep_alive: KeepAliveValue – Model persistence control

PIIDeidentifyRequest

The PIIDeidentifyRequest schema (lines 172-186) defines the contract for de-identification operations including masking, removal, replacement, hashing, and date shifting.

Key fields:

  • text: str – Source document
  • method: Literal["mask","remove","replace","hash","shift_dates"] – De-identification strategy, defaults to "mask"
  • model_name: str – Defaults to _DEFAULT_PII_MODEL
  • confidence_threshold: float – Higher default of 0.7 for de-identification safety
  • keep_year: bool – Defaults to True
  • shift_dates: Optional[bool] – Enable date shifting
  • date_shift_days: Optional[int] – Days to shift dates
  • keep_mapping: bool – Defaults to False
  • use_smart_merging: bool – Defaults to True
  • use_safety_sweep: bool – Defaults to True
  • lang: PIILanguage – Defaults to "en"
  • normalize_accents: Optional[bool]
  • keep_alive: KeepAliveValue

This model includes cross-field validation via _validate_shift_dates to ensure shift_dates and date_shift_days are used consistently.

ModelUnloadRequest

The ModelUnloadRequest schema (lines 219-227) manages model unloading from the server.

Key fields:

  • model_name: Optional[str] – Target model identifier, validated via validate_model_name if provided
  • all: bool – Defaults to False

Validation rule: If all is False and model_name is omitted, the schema raises a validation error: "model_name is required unless all=true".

Common Schema Building Blocks

All request models inherit from _StrictModel, defined in openmed/service/schemas.py (lines 91-100). This base class configures model_config = ConfigDict(extra="forbid"), ensuring that any unexpected fields in the payload trigger immediate validation errors.

Shared type definitions:

  • _DEFAULT_PII_MODEL – Central constant defined as "OpenMed/OpenMed-PII-SuperClinical-Small-44M-v1" (lines 28-29)
  • PIILanguage – Literal union of 12 supported languages (lines 37-39), synchronized with openmed.core.pii_i18n.SUPPORTED_LANGUAGES

Validation and Field Processing

The schema layer implements several validation helpers that execute during model instantiation:

  • _normalize_text – Trims whitespace, checks non-emptiness, and enforces maximum character limits via get_max_text_length from openmed/service/limits.py
  • _normalize_model_name – Forwards to validate_model_name in openmed/utils/validation.py
  • _normalize_confidence_threshold – Forwards to validate_confidence_threshold to clamp values between 0 and 1
  • _validate_keep_alive_value – Parses keep-alive specifications using parse_keep_alive from openmed/service/keep_alive.py
  • _normalize_shift_dates_payload – Reconciles shift_dates, method, and date_shift_days fields for consistent de-identification requests

Working with the Schemas

Building an Analysis Request

from openmed.service.schemas import AnalyzeRequest

req = AnalyzeRequest(
    text="Patient presents with fever and cough.",
    model_name="disease_detection_superclinical",
    confidence_threshold=0.5,
    keep_alive=300,  # keep the model warm for 5 minutes

)

print(req.json())

Extracting PII with Language Selection

from openmed.service.schemas import PIIExtractRequest

req = PIIExtractRequest(
    text="John Doe, 123-45-6789, visited on 2023-02-01.",
    lang="en",  # must be one of the PIILanguage literals

    keep_alive="auto",  # let the server decide keep-alive duration

)

print(req.json())

De-identifying with Date Shifting

from openmed.service.schemas import PIIDeidentifyRequest

req = PIIDeidentifyRequest(
    text="Patient admitted on 2022-12-15.",
    method="shift_dates",
    date_shift_days=30,  # shift all dates forward by 30 days

    keep_alive=60,
)

print(req.json())

Unloading a Specific Model

from openmed.service.schemas import ModelUnloadRequest

req = ModelUnloadRequest(model_name="disease_detection_superclinical")
print(req.json())

Summary

  • OpenMed's database schema is implemented as Pydantic request models rather than SQL tables, defined in openmed/service/schemas.py.
  • Four primary schemas govern all data input: AnalyzeRequest, PIIExtractRequest, PIIDeidentifyRequest, and ModelUnloadRequest.
  • Strict validation occurs through _StrictModel (extra="forbid"), field-specific validators, and cross-field validation methods.
  • Default values are explicitly typed: confidence thresholds range from 0.0 to 0.7 depending on the endpoint, and PII operations default to the SuperClinical-Small-44M model.
  • Language support for PII operations is restricted to 12 literal values defined in PIILanguage and synchronized with openmed/core/pii_i18n.py.

Frequently Asked Questions

Does OpenMed use a traditional SQL database schema?

No. According to the OpenMed source code, the system does not expose a relational database schema. Instead, it uses Pydantic request schemas in openmed/service/schemas.py to define data structures, enforce types, and validate payloads before processing. These Python classes act as the contract between client and server.

What validation rules apply to the confidence_threshold field?

The confidence_threshold field is clamped to a 0-1 range via validate_confidence_threshold in openmed/utils/validation.py. In AnalyzeRequest it defaults to 0.0, in PIIExtractRequest to 0.5, and in PIIDeidentifyRequest to 0.7 for higher safety margins during de-identification.

How do I specify language support in PII requests?

Use the lang field with a literal value from the PIILanguage type definition. Valid options are: "en", "fr", "de", "it", "es", "nl", "hi", "te", "pt", "ar", "ja", and "tr". This list is maintained in openmed/core/pii_i18n.py and referenced in openmed/service/schemas.py (lines 37-39).

What happens if I send extra fields in a request payload?

The request will fail validation immediately. All OpenMed request models inherit from _StrictModel, which configures extra="forbid" in its Pydantic model_config. This setting ensures that any unexpected fields trigger a validation error rather than being silently ignored.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →