How OpenMed's Multi-Backend Architecture Handles Model Routing

OpenMed's multi-backend architecture routes inference requests through a central get_backend() function that auto-detects the optimal execution engine—prioritizing Apple MLX on Silicon hardware while falling back to HuggingFace Transformers—and returns a standardized pipeline interface that produces HuggingFace-compatible entity dictionaries.

OpenMed separates model loading from inference execution through a robust backend abstraction layer defined in openmed/core/backends.py. This design enables seamless hardware optimization by automatically selecting between Apple's MLX framework for Apple Silicon and standard PyTorch-based pipelines for other platforms. Understanding how OpenMed's multi-backend architecture handles model routing is essential for deploying medical NLP models across diverse environments without modifying application code.

Central Routing Logic

The get_backend() function (lines 15-48 in openmed/core/backends.py) serves as the single entry point for all backend selection decisions. It implements a detection hierarchy that balances performance optimization with deployment flexibility.

Explicit Backend Selection

When callers specify a preferred engine via the name parameter (e.g., name="hf" or name="mlx"), the router validates the string, instantiates the corresponding HuggingFaceBackend or MLXBackend class, and immediately raises a RuntimeError if the requested dependencies are missing.

from openmed.core.backends import get_backend

# Force HuggingFace backend; raises if transformers not installed

hf_backend = get_backend(name="hf", config=my_config)
pipeline = hf_backend.create_pipeline("OpenMed/OpenMed-NER-PharmaDetect-SuperClinical-434M")

Source: explicit-name handling block (lines 29-40)

Auto-Detection Priority

When name=None, the function iterates over the priority tuple ("mlx", "hf"), constructing each backend candidate and invoking is_available() until it locates a workable implementation. The MLX backend validates Apple Silicon hardware and the presence of mlx.core, while the HuggingFace backend verifies transformers installation. If neither backend reports availability, the function raises a RuntimeError with installation guidance.

from openmed.core.backends import get_backend

# Automatically selects MLX on Apple Silicon, otherwise HuggingFace

backend = get_backend(config=my_config)  # name defaults to None

pipeline = backend.create_pipeline("OpenMed/OpenMed-NER-DiseaseDetect-SuperClinical-434M")
result = pipeline("Patient shows signs of hypertension.")

Source: auto-detect loop (lines 42-48)

Privacy-Filter Routing Specialization

Privacy-filter models implement additional routing logic in openmed/core/backends.py because certain artifacts are MLX-only. The system uses specialized helper functions to handle hardware-specific variants without exposing complexity to the caller.

Backend Selection Logic

The select_privacy_filter_backend() function (lines 87-99) determines the appropriate engine by checking:

  • Whether "mlx" appears in the model name
  • If the artifact path indicates an MLX-only model via _is_privacy_filter_artifact_path
  • The return value of MLXBackend().is_available()

Automatic Model Substitution

For non-Apple-Silicon hosts, resolve_privacy_filter_model() (lines 101-119) intercepts MLX-only model requests and substitutes the equivalent Torch-based variant. This function emits a one-time UserWarning upon the first substitution to inform developers of the automatic fallback behavior.

The create_privacy_filter_pipeline() convenience function (lines 121-139) orchestrates these helpers to return a backend-agnostic pipeline:

from openmed.core.backends import create_privacy_filter_pipeline

# On Linux, automatically falls back to Torch equivalent

pipeline = create_privacy_filter_pipeline("OpenMed/privacy-filter-mlx")
entities = pipeline("John Doe's SSN is 123-45-6789.")

Source: privacy-filter routing implementation

Backend Implementation Interfaces

Each backend conforms to the InferenceBackend protocol, ensuring consistent behavior while encapsulating framework-specific optimizations.

HuggingFaceBackend

The HuggingFaceBackend class utilizes ModelLoader._create_hf_pipeline() (defined in openmed/core/models.py) to construct PyTorch-based token classification pipelines. It returns entity dictionaries matching the HuggingFace token-classification format: [{ "entity_group": ..., "score": ..., "word": ..., "start": ..., "end": ... }].

MLXBackend

The MLXBackend class routes to openmed.mlx.inference.create_mlx_pipeline() (lines 106-112 in openmed/mlx/inference.py) to build Apple-optimized pipelines. Despite the different underlying framework, it returns identically structured results, ensuring downstream components (entity merging, quality gates, output formatters) remain unchanged regardless of the execution engine.

from openmed.mlx.inference import create_mlx_pipeline

# Direct MLX pipeline construction when backend is known

mlx_pipe = create_mlx_pipeline(
    model_name="OpenMed/OpenMed-NER-DiseaseDetect-SuperClinical-434M",
    aggregation_strategy="simple",
)
print(mlx_pipe("The biopsy revealed carcinoma."))

Summary

  • Separation of concerns: OpenMed's architecture isolates model loading (openmed/core/backends.py) from inference execution, allowing dynamic backend selection at runtime.
  • Priority-based auto-detection: The get_backend() function prioritizes MLX over HuggingFace when name=None, maximizing performance on Apple Silicon while maintaining cross-platform compatibility.
  • Unified interface: All backends return HuggingFace-compatible entity dictionaries, ensuring that application code remains backend-agnostic.
  • Privacy-filter intelligence: Specialized routing in select_privacy_filter_backend() and resolve_privacy_filter_model() handles MLX-only artifacts and automatic Torch fallbacks without manual intervention.
  • Explicit control: Developers can force specific backends via the name parameter, with clear error messages when dependencies are missing.

Frequently Asked Questions

How does OpenMed decide which backend to use?

When name=None, OpenMed iterates through the tuple ("mlx", "hf") and selects the first backend where is_available() returns True. This prioritizes MLX on Apple Silicon devices and automatically falls back to HuggingFace Transformers on other hardware.

What happens if I specify a backend that isn't installed?

If you explicitly request "hf" or "mlx" via the name parameter and the required dependencies are missing, get_backend() raises a RuntimeError with specific installation instructions for the requested framework.

Can privacy-filter models run on non-Apple hardware?

Yes. The resolve_privacy_filter_model() function automatically detects non-Apple-Silicon environments and substitutes MLX-only privacy-filter models with their Torch equivalents, emitting a UserWarning to notify you of the substitution.

What data format do pipelines return regardless of the backend?

All OpenMed pipelines return a list of dictionaries following the HuggingFace token-classification schema: [{ "entity_group": str, "score": float, "word": str, "start": int, "end": int }]. This standardization ensures that entity merging and quality-gate logic work identically across MLX and HuggingFace backends.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →