How to Select a Model from OpenMed's Registry: A Complete Guide
You can select an OpenMed model by querying the built-in registry using the ModelLoader class, which maps symbolic keys like disease_detection_tiny to full HuggingFace model IDs via the OPENMED_MODELS dictionary defined in openmed/core/model_registry.py.
Selecting the right medical NLP model from OpenMed's registry allows you to leverage pre-trained biomedical models without hard-coding HuggingFace URLs. The maziyarpanahi/openmed repository provides a centralized model registry that loads metadata from a static manifest and exposes convenient lookup methods for filtering by category, size, or clinical domain.
How the OpenMed Model Registry Is Constructed
The registry system separates model discovery from model loading. At import time, openmed/core/model_registry.py builds a module-level constant OPENMED_MODELS by parsing the static models.jsonl manifest file located at the repository root.
Loading the Manifest and Creating ModelInfo Objects
The load_manifest_rows() function reads the JSONL file (defined at MANIFEST_PATH = Path(__file__).resolve().parents[2] / "models.jsonl"). Each row is converted into a ModelInfo dataclass by _model_info_from_row(), which extracts fields including category, display name, entity types, size category, and recommended confidence thresholds.
Generating Registry Keys
The _registry_key() function categorizes models into families—Privacy (PII), NER, or general—while _unique_key() appends format-specific suffixes to ensure uniqueness when collisions occur. Finally, _compatibility_aliases() adds legacy aliases to maintain backward compatibility with older codebases.
Querying the Registry to Select Models
Once loaded, the registry exposes several helper functions to locate specific models without manually searching the JSONL file:
get_model_info(key)– Returns aModelInfodataclass for a registry key or full repo ID.get_models_by_category(cat)– Filters models by clinical category (e.g., Disease, Oncology, Privacy).get_models_by_size(size)– Returns models matching size buckets like Tiny, Small, or Base.get_model_suggestions(text)– Heuristically suggests up to three models based on keywords in clinical text.list_model_categories()– Returns all available categories present in the registry.get_pii_models_by_language(lang)– Retrieves language-specific PII models (e.g.,pii_fr_...).
Listing All Available Models
To view every registry key programmatically:
from openmed.core.models import ModelLoader
loader = ModelLoader()
print(loader.list_available_models())
Retrieving Detailed Model Metadata
Access the ModelInfo dataclass for a specific key to inspect its HuggingFace ID, entity types, and recommended confidence:
from openmed.core.models import ModelLoader
loader = ModelLoader()
info = loader.get_registry_info('oncology_detection_tiny')
print(info)
Finding Models by Clinical Context
The registry can suggest appropriate models based on raw clinical text heuristics:
from openmed.core.models import ModelLoader
loader = ModelLoader()
text = "Patient was diagnosed with metastatic breast cancer and started on paclitaxel."
suggestions = loader.get_model_suggestions(text)
for key, model, reason in suggestions:
print(f"{key}: {model.display_name} ({reason})")
Filtering PII Models by Language
For privacy-focused applications, select language-specific models using the registry's dedicated helper:
from openmed.core.model_registry import get_pii_models_by_language
fr_models = get_pii_models_by_language('fr')
for key, info in fr_models.items():
print(key, info.display_name)
Loading and Running Models from the Registry
The ModelLoader class in openmed/core/models.py bridges the registry and HuggingFace's infrastructure. When you request a model, the loader uses _resolve_model_name() to translate registry keys to full model IDs, then caches the model, tokenizer, and pipeline for reuse.
Creating an Inference Pipeline
from openmed.core.models import ModelLoader
loader = ModelLoader()
pipeline = loader.create_pipeline('disease_detection_tiny')
result = pipeline("The patient suffers from diabetes mellitus.")
print(result)
Summary
- The OpenMed registry is defined in
openmed/core/model_registry.pyand loads frommodels.jsonlat runtime to populate theOPENMED_MODELSdictionary. - Use
ModelLoaderto query available models without hard-coding HuggingFace URLs, leveraging methods likelist_available_models()andget_registry_info(). - Filter models by clinical category, size bucket, or language using helpers such as
get_models_by_category()andget_pii_models_by_language(). - The
get_model_suggestions()function analyzes clinical text to recommend relevant models based on detected keywords. ModelLoadercaches loaded models and tokenizers to avoid redundant downloads and API calls.
Frequently Asked Questions
How do I find the registry key for a specific OpenMed model?
You can call loader.list_available_models() on a ModelLoader instance to see all available keys, or use get_model_suggestions() with sample text from your clinical domain to receive heuristic recommendations. Each key maps to a ModelInfo object containing the full HuggingFace model ID as the model_id attribute.
What is the difference between the registry key and the model ID?
The registry key is a human-readable shorthand (e.g., disease_detection_tiny) used within OpenMed's ecosystem, while the model ID is the full HuggingFace repository path (e.g., OpenMed/OpenMed-NER-DiseaseDetection-TinyMed-65M). The ModelLoader automatically translates keys to IDs using the internal _resolve_model_name() method.
Can I filter models by supported languages?
Yes. For PII detection models, use get_pii_models_by_language(lang) from openmed/core/model_registry.py, passing the ISO language code (e.g., 'fr' for French). For other categories, inspect the languages field of the ModelInfo dataclass returned by get_model_info().
Where is the model metadata stored?
The static metadata lives in models.jsonl at the repository root, which is parsed at import time by load_manifest_rows() in openmed/core/model_registry.py. This file contains the definitive list of supported models, their categories, sizes, supported entity types, and compatibility aliases.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →