How to Select a Model from OpenMed's Registry: A Complete Guide

You can select an OpenMed model by querying the built-in registry using the ModelLoader class, which maps symbolic keys like disease_detection_tiny to full HuggingFace model IDs via the OPENMED_MODELS dictionary defined in openmed/core/model_registry.py.

Selecting the right medical NLP model from OpenMed's registry allows you to leverage pre-trained biomedical models without hard-coding HuggingFace URLs. The maziyarpanahi/openmed repository provides a centralized model registry that loads metadata from a static manifest and exposes convenient lookup methods for filtering by category, size, or clinical domain.

How the OpenMed Model Registry Is Constructed

The registry system separates model discovery from model loading. At import time, openmed/core/model_registry.py builds a module-level constant OPENMED_MODELS by parsing the static models.jsonl manifest file located at the repository root.

Loading the Manifest and Creating ModelInfo Objects

The load_manifest_rows() function reads the JSONL file (defined at MANIFEST_PATH = Path(__file__).resolve().parents[2] / "models.jsonl"). Each row is converted into a ModelInfo dataclass by _model_info_from_row(), which extracts fields including category, display name, entity types, size category, and recommended confidence thresholds.

Generating Registry Keys

The _registry_key() function categorizes models into families—Privacy (PII), NER, or general—while _unique_key() appends format-specific suffixes to ensure uniqueness when collisions occur. Finally, _compatibility_aliases() adds legacy aliases to maintain backward compatibility with older codebases.

Querying the Registry to Select Models

Once loaded, the registry exposes several helper functions to locate specific models without manually searching the JSONL file:

  • get_model_info(key) – Returns a ModelInfo dataclass for a registry key or full repo ID.
  • get_models_by_category(cat) – Filters models by clinical category (e.g., Disease, Oncology, Privacy).
  • get_models_by_size(size) – Returns models matching size buckets like Tiny, Small, or Base.
  • get_model_suggestions(text) – Heuristically suggests up to three models based on keywords in clinical text.
  • list_model_categories() – Returns all available categories present in the registry.
  • get_pii_models_by_language(lang) – Retrieves language-specific PII models (e.g., pii_fr_...).

Listing All Available Models

To view every registry key programmatically:

from openmed.core.models import ModelLoader

loader = ModelLoader()
print(loader.list_available_models())

Retrieving Detailed Model Metadata

Access the ModelInfo dataclass for a specific key to inspect its HuggingFace ID, entity types, and recommended confidence:

from openmed.core.models import ModelLoader

loader = ModelLoader()
info = loader.get_registry_info('oncology_detection_tiny')
print(info)

Finding Models by Clinical Context

The registry can suggest appropriate models based on raw clinical text heuristics:

from openmed.core.models import ModelLoader

loader = ModelLoader()
text = "Patient was diagnosed with metastatic breast cancer and started on paclitaxel."
suggestions = loader.get_model_suggestions(text)
for key, model, reason in suggestions:
    print(f"{key}: {model.display_name} ({reason})")

Filtering PII Models by Language

For privacy-focused applications, select language-specific models using the registry's dedicated helper:

from openmed.core.model_registry import get_pii_models_by_language

fr_models = get_pii_models_by_language('fr')
for key, info in fr_models.items():
    print(key, info.display_name)

Loading and Running Models from the Registry

The ModelLoader class in openmed/core/models.py bridges the registry and HuggingFace's infrastructure. When you request a model, the loader uses _resolve_model_name() to translate registry keys to full model IDs, then caches the model, tokenizer, and pipeline for reuse.

Creating an Inference Pipeline

from openmed.core.models import ModelLoader

loader = ModelLoader()
pipeline = loader.create_pipeline('disease_detection_tiny')
result = pipeline("The patient suffers from diabetes mellitus.")
print(result)

Summary

  • The OpenMed registry is defined in openmed/core/model_registry.py and loads from models.jsonl at runtime to populate the OPENMED_MODELS dictionary.
  • Use ModelLoader to query available models without hard-coding HuggingFace URLs, leveraging methods like list_available_models() and get_registry_info().
  • Filter models by clinical category, size bucket, or language using helpers such as get_models_by_category() and get_pii_models_by_language().
  • The get_model_suggestions() function analyzes clinical text to recommend relevant models based on detected keywords.
  • ModelLoader caches loaded models and tokenizers to avoid redundant downloads and API calls.

Frequently Asked Questions

How do I find the registry key for a specific OpenMed model?

You can call loader.list_available_models() on a ModelLoader instance to see all available keys, or use get_model_suggestions() with sample text from your clinical domain to receive heuristic recommendations. Each key maps to a ModelInfo object containing the full HuggingFace model ID as the model_id attribute.

What is the difference between the registry key and the model ID?

The registry key is a human-readable shorthand (e.g., disease_detection_tiny) used within OpenMed's ecosystem, while the model ID is the full HuggingFace repository path (e.g., OpenMed/OpenMed-NER-DiseaseDetection-TinyMed-65M). The ModelLoader automatically translates keys to IDs using the internal _resolve_model_name() method.

Can I filter models by supported languages?

Yes. For PII detection models, use get_pii_models_by_language(lang) from openmed/core/model_registry.py, passing the ISO language code (e.g., 'fr' for French). For other categories, inspect the languages field of the ModelInfo dataclass returned by get_model_info().

Where is the model metadata stored?

The static metadata lives in models.jsonl at the repository root, which is parsed at import time by load_manifest_rows() in openmed/core/model_registry.py. This file contains the definitive list of supported models, their categories, sizes, supported entity types, and compatibility aliases.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →