# How OpenMed's Multi-Backend Architecture Handles Model Routing

> OpenMed's multi-backend architecture smartly routes inference requests. Discover how it auto-detects and prioritizes MLX on Silicon, falling back to HuggingFace Transformers for optimal performance.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: architecture
- Published: 2026-06-10

---

**OpenMed's multi-backend architecture routes inference requests through a central `get_backend()` function that auto-detects the optimal execution engine—prioritizing Apple MLX on Silicon hardware while falling back to HuggingFace Transformers—and returns a standardized pipeline interface that produces HuggingFace-compatible entity dictionaries.**

OpenMed separates model loading from inference execution through a robust backend abstraction layer defined in [`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py). This design enables seamless hardware optimization by automatically selecting between Apple's MLX framework for Apple Silicon and standard PyTorch-based pipelines for other platforms. Understanding how OpenMed's multi-backend architecture handles model routing is essential for deploying medical NLP models across diverse environments without modifying application code.

## Central Routing Logic

The `get_backend()` function (lines 15-48 in [`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py)) serves as the single entry point for all backend selection decisions. It implements a detection hierarchy that balances performance optimization with deployment flexibility.

### Explicit Backend Selection

When callers specify a preferred engine via the `name` parameter (e.g., `name="hf"` or `name="mlx"`), the router validates the string, instantiates the corresponding `HuggingFaceBackend` or `MLXBackend` class, and immediately raises a `RuntimeError` if the requested dependencies are missing.

```python
from openmed.core.backends import get_backend

# Force HuggingFace backend; raises if transformers not installed

hf_backend = get_backend(name="hf", config=my_config)
pipeline = hf_backend.create_pipeline("OpenMed/OpenMed-NER-PharmaDetect-SuperClinical-434M")

```

*Source*: explicit-name handling block (lines 29-40)

### Auto-Detection Priority

When `name=None`, the function iterates over the priority tuple `("mlx", "hf")`, constructing each backend candidate and invoking `is_available()` until it locates a workable implementation. The **MLX backend** validates Apple Silicon hardware and the presence of `mlx.core`, while the **HuggingFace backend** verifies `transformers` installation. If neither backend reports availability, the function raises a `RuntimeError` with installation guidance.

```python
from openmed.core.backends import get_backend

# Automatically selects MLX on Apple Silicon, otherwise HuggingFace

backend = get_backend(config=my_config)  # name defaults to None

pipeline = backend.create_pipeline("OpenMed/OpenMed-NER-DiseaseDetect-SuperClinical-434M")
result = pipeline("Patient shows signs of hypertension.")

```

*Source*: auto-detect loop (lines 42-48)

## Privacy-Filter Routing Specialization

Privacy-filter models implement additional routing logic in [`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py) because certain artifacts are **MLX-only**. The system uses specialized helper functions to handle hardware-specific variants without exposing complexity to the caller.

### Backend Selection Logic

The `select_privacy_filter_backend()` function (lines 87-99) determines the appropriate engine by checking:
- Whether "mlx" appears in the model name
- If the artifact path indicates an MLX-only model via `_is_privacy_filter_artifact_path`
- The return value of `MLXBackend().is_available()`

### Automatic Model Substitution

For non-Apple-Silicon hosts, `resolve_privacy_filter_model()` (lines 101-119) intercepts MLX-only model requests and substitutes the equivalent Torch-based variant. This function emits a one-time `UserWarning` upon the first substitution to inform developers of the automatic fallback behavior.

The `create_privacy_filter_pipeline()` convenience function (lines 121-139) orchestrates these helpers to return a backend-agnostic pipeline:

```python
from openmed.core.backends import create_privacy_filter_pipeline

# On Linux, automatically falls back to Torch equivalent

pipeline = create_privacy_filter_pipeline("OpenMed/privacy-filter-mlx")
entities = pipeline("John Doe's SSN is 123-45-6789.")

```

*Source*: privacy-filter routing implementation

## Backend Implementation Interfaces

Each backend conforms to the `InferenceBackend` protocol, ensuring consistent behavior while encapsulating framework-specific optimizations.

### HuggingFaceBackend

The `HuggingFaceBackend` class utilizes `ModelLoader._create_hf_pipeline()` (defined in [`openmed/core/models.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/models.py)) to construct PyTorch-based token classification pipelines. It returns entity dictionaries matching the HuggingFace token-classification format: `[{ "entity_group": ..., "score": ..., "word": ..., "start": ..., "end": ... }]`.

### MLXBackend

The `MLXBackend` class routes to `openmed.mlx.inference.create_mlx_pipeline()` (lines 106-112 in [`openmed/mlx/inference.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/inference.py)) to build Apple-optimized pipelines. Despite the different underlying framework, it returns identically structured results, ensuring downstream components (entity merging, quality gates, output formatters) remain unchanged regardless of the execution engine.

```python
from openmed.mlx.inference import create_mlx_pipeline

# Direct MLX pipeline construction when backend is known

mlx_pipe = create_mlx_pipeline(
    model_name="OpenMed/OpenMed-NER-DiseaseDetect-SuperClinical-434M",
    aggregation_strategy="simple",
)
print(mlx_pipe("The biopsy revealed carcinoma."))

```

## Summary

- **Separation of concerns**: OpenMed's architecture isolates model loading ([`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py)) from inference execution, allowing dynamic backend selection at runtime.
- **Priority-based auto-detection**: The `get_backend()` function prioritizes MLX over HuggingFace when `name=None`, maximizing performance on Apple Silicon while maintaining cross-platform compatibility.
- **Unified interface**: All backends return HuggingFace-compatible entity dictionaries, ensuring that application code remains backend-agnostic.
- **Privacy-filter intelligence**: Specialized routing in `select_privacy_filter_backend()` and `resolve_privacy_filter_model()` handles MLX-only artifacts and automatic Torch fallbacks without manual intervention.
- **Explicit control**: Developers can force specific backends via the `name` parameter, with clear error messages when dependencies are missing.

## Frequently Asked Questions

### How does OpenMed decide which backend to use?

When `name=None`, OpenMed iterates through the tuple `("mlx", "hf")` and selects the first backend where `is_available()` returns `True`. This prioritizes MLX on Apple Silicon devices and automatically falls back to HuggingFace Transformers on other hardware.

### What happens if I specify a backend that isn't installed?

If you explicitly request `"hf"` or `"mlx"` via the `name` parameter and the required dependencies are missing, `get_backend()` raises a `RuntimeError` with specific installation instructions for the requested framework.

### Can privacy-filter models run on non-Apple hardware?

Yes. The `resolve_privacy_filter_model()` function automatically detects non-Apple-Silicon environments and substitutes MLX-only privacy-filter models with their Torch equivalents, emitting a `UserWarning` to notify you of the substitution.

### What data format do pipelines return regardless of the backend?

All OpenMed pipelines return a list of dictionaries following the HuggingFace token-classification schema: `[{ "entity_group": str, "score": float, "word": str, "start": int, "end": int }]`. This standardization ensures that entity merging and quality-gate logic work identically across MLX and HuggingFace backends.