How to Contribute to OpenMed: A Complete Guide for Healthcare AI Developers
To contribute to OpenMed, fork the repository, create a feature branch, add your code with accompanying tests in tests/unit/, run make test to verify locally, and submit a pull request for automated CI validation.
OpenMed is a local‑first healthcare AI library developed by maziyarpanahi/openmed that runs entirely on‑device for privacy‑sensitive NER and PII extraction tasks. The repository is structured to welcome contributors who want to add new models, improve core pipelines, or extend documentation without requiring special access or secrets. This guide walks through the exact file paths, architectural patterns, and submission workflow used by the maintainers.
Repository Structure and Key Entry Points
Understanding the codebase layout is essential before making changes. The project follows a modular architecture that separates core logic, backend implementations, and user interfaces.
Core Library Components
The heart of OpenMed lives in openmed/core/, where the NER pipelines, PII filters, and model registry logic reside. Key files include:
openmed/core/models.py– Defines model loading and inference abstractionsopenmed/core/model_registry.py– Manages themodels.jsonlcatalog that registers available model familiesopenmed/core/backends.py– Contains theBackendinterface that abstracts PyTorch, CUDA, and Apple MLX implementations
Backend Implementations
Hardware‑specific optimizations are isolated in openmed/torch/ and future backend directories. For example, openmed/torch/privacy_filter.py implements PyTorch‑specific PII filtering logic. Adding a new inference engine (such as a Rust‑based runtime) only requires implementing the Backend interface in openmed/core/backends.py and registering it in the appropriate module.
Service and CLI Layers
openmed/service/app.py– FastAPI application exposing/analyzeand/pii/extractendpointsopenmed/cli/main.py– Typer‑based command‑line interface entry pointdocs/– MkDocs source files, including the detaileddocs/contributing.mdguide
Step‑by‑Step Contribution Workflow
Follow this exact sequence to ensure your changes integrate smoothly with the CI pipeline:
-
Fork and clone the
maziyarpanahi/openmedrepository to your GitHub account and local machine. -
Create a branch using a descriptive naming convention such as
feat/awesome-new-modelorfix/ner-tokenization-bug. -
Add or modify code in the appropriate directory, keeping changes scoped to a single feature or bug fix. For model additions, update
openmed/core/model_registry.pyand themodels.jsonlcatalog. -
Add tests in
tests/unit/that exercise the new behavior. The test suite uses pytest and must maintain 100% pass rate for core functionality. -
Run the local test suite using the provided Makefile shortcuts:
make test # Runs pytest with coverage reporting make docs-serve # Optional: preview documentation locallyThese targets are defined in the root
Makefile, which also includes helpers for versioning and building. -
Open a Pull Request. CI automatically executes the full test matrix across Linux and macOS, testing PyTorch + CUDA and MLX backends, and validates the MkDocs build with
mkdocs build --strict. -
Iterate on review feedback. The maintainers prefer small, focused PRs that address a single concern as outlined in
docs/contributing.md. -
Release process (maintainers only): Upon merge, version bumps occur via
make bump-patch(or minor/major), wheels build viamake build, and git tags trigger the.github/workflows/publish.ymlworkflow to publish to PyPI and update GitHub Pages.
Practical Contribution Examples
Adding a New NER Model
To integrate a custom model stored locally without network calls:
from openmed import OpenMedConfig, analyze_text
# Configure for CPU inference
config = OpenMedConfig(device="cpu")
# Analyze text using a local model path
result = analyze_text(
"Patient presents with hypertension and diabetes.",
model_id="./my_models/disease_detection_custom",
config=config,
)
# Process entities
for ent in result.entities:
print(f"{ent.label:<12} {ent.text:<30} {ent.confidence:.2f}")
The model_id parameter bypasses the Hugging Face Hub, which is essential for on‑premise healthcare deployments. Update openmed/core/model_registry.py if adding the model to the permanent registry.
Extending the CLI
Add new sub‑commands by modifying openmed/cli/main.py:
import typer
from openmed import BatchProcessor
app = typer.Typer()
@app.command()
def batch_process(
model: str,
input_file: str,
output_file: str,
batch_size: int = 16,
):
"""Run a batch of texts through a model and write JSON results."""
bp = BatchProcessor(model_name=model, batch_size=batch_size)
with open(input_file) as f:
texts = f.read().splitlines()
results = bp.process_texts(texts)
with open(output_file, "w") as f:
import json
json.dump(results, f, indent=2)
if __name__ == "__main__":
app()
Test your new command locally:
openmed batch-process \
--model disease_detection_superclinical \
--input-file docs/example_texts.txt \
--output-file results.json
Writing Unit Tests
Create a new file in tests/unit/ following the existing patterns:
# tests/unit/test_my_new_feature.py
from openmed import analyze_text
def test_custom_model_returns_entities():
"""Verify that custom models correctly identify disease entities."""
result = analyze_text(
"Patient has asthma.",
model_name="disease_detection_superclinical",
)
assert any(
ent.label == "DISEASE" and "asthma" in ent.text.lower()
for ent in result.entities
)
Run make test to execute your new test alongside the existing suite covering core NER, PII extraction, multilingual handling, and the REST service.
Understanding the Architecture for Contributors
Modular backends allow hardware‑specific optimizations without touching core logic. The Backend interface in openmed/core/backends.py abstracts inference engines, enabling contributors to add CUDA optimizations or new frameworks like MLX without breaking existing code.
Model registry uses a JSON‑line catalog (models.jsonl) managed by openmed/core/model_registry.py. New model families require entries in this catalog and corresponding loader implementations.
Testing pipeline enforces quality through comprehensive unit tests in tests/unit/test_core.py and integration tests. Any change must maintain backwards compatibility and keep the test suite green, as CI runs the full matrix on every PR.
Documentation‑first development requires that new features include Markdown guides under docs/. The MkDocs CI validates that the site builds without warnings using mkdocs build --strict.
Summary
- Fork and branch the
maziyarpanahi/openmedrepository before making changes - Target specific directories:
openmed/core/for pipeline logic,openmed/torch/for backend code,openmed/cli/for commands, andtests/unit/for validation - Use
make testto verify changes locally before submitting - Update
docs/when adding features to ensure the MkDocs site builds correctly - Submit focused PRs that pass the automated CI matrix testing PyTorch, CUDA, and MLX backends
Frequently Asked Questions
What areas of OpenMed need contributors most?
The project actively seeks contributions for new NER models (adding entries to models.jsonl and openmed/core/model_registry.py), backend optimizations (extending openmed/core/backends.py for additional hardware), and documentation improvements (guides under docs/). The REST service in openmed/service/app.py and CLI in openmed/cli/main.py also welcome feature extensions.
Do I need special access or API keys to contribute?
No. OpenMed is designed as a local‑first library that operates entirely on‑device. Contributors can fork the repository, install development dependencies via pip install -e ".[dev]" as defined in pyproject.toml, and run the full test suite without API keys or cloud access.
How do I test my changes locally before submitting a PR?
Run make test from the repository root to execute pytest with coverage reporting. For documentation changes, run make docs-serve to preview the MkDocs site locally. The Makefile provides additional shortcuts like make build for creating wheels and make bump-patch for version management.
What is the release process for OpenMed?
Only maintainers execute releases. After merging a PR, maintainers run make bump-patch (or minor/major) to update version numbers, then make build to generate wheels. Pushing a git tag (vX.Y.Z) triggers the .github/workflows/publish.yml GitHub Actions workflow, which publishes the package to PyPI and updates the GitHub Pages documentation site.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →