# Setting up MLX Backend on Apple Silicon for Accelerated Inference with OpenMed

> Accelerate OpenMed inference on Apple Silicon with the MLX backend. Achieve 24-33x speedups over CPU PyTorch automatically. No code changes needed for seamless integration.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-12

---

**OpenMed automatically utilizes the Apple MLX backend on Apple Silicon devices, delivering 24-33× speed-ups over CPU-only PyTorch inference without requiring any code changes.**

OpenMed provides a pluggable inference-backend abstraction that automatically selects the fastest available runtime for medical NLP workloads. When running on macOS, iPadOS, or iOS with Apple Silicon chips, the library prioritizes the Apple MLX framework for token classification and privacy filtering tasks. This guide explains how to configure and verify the MLX backend for accelerated inference using the OpenMed source code.

## How OpenMed Auto-Detects the MLX Backend

The backend selection logic lives in [`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py). The `get_backend()` function queries `MLXBackend.is_available()` to determine whether the host supports MLX acceleration. This check returns `True` only when the platform is `Darwin` (macOS/iOS) and the `mlx` Python package can be imported successfully.

When these conditions are met, `get_backend()` automatically selects the MLX backend; otherwise, it falls back to the HuggingFace/PyTorch backend. This auto-detection happens at runtime, ensuring that scripts written on Apple Silicon will seamlessly transition to CPU-based inference on Linux or Windows hosts without modification.

```python
from openmed.core.backends import get_backend

# Auto-detect: returns MLXBackend on Apple Silicon, PyTorchBackend elsewhere

backend = get_backend()

# Force specific backend (raises RuntimeError if MLX unavailable)

backend = get_backend(name="mlx")

```

## Installing MLX Dependencies

MLX support is shipped as an optional extra to keep base installations lightweight. The dependency specification resides in [`pyproject.toml`](https://github.com/maziyarpanahi/openmed/blob/main/pyproject.toml) under `[project.optional-dependencies]`, where the `mlx` extra pulls in `mlx`, `transformers`, `tokenizers`, `safetensors`, and `huggingface-hub`.

Install the accelerated backend on your Apple Silicon machine:

```bash

# From PyPI

pip install "openmed[mlx]"

# Or from source

pip install -e ".[mlx]"

```

Verify the installation by confirming that MLX is importable:

```bash
python -c "import mlx.core; print('MLX available')"

```

## Using the MLX Backend

### Automatic Backend Selection

Once installed, the standard OpenMed API automatically routes inference through the MLX pipeline. The `analyze_text()` function utilizes the auto-detected backend without additional configuration.

```python
from openmed import analyze_text

result = analyze_text(
    "Patient John Doe, DOB 1990-05-15, SSN 123-45-6789",
    model_name="pii_superclinical_large",   # Any MLX-compatible model

)
print(result.entities)   # Uses MLX acceleration on Apple Silicon

```

### Explicit Configuration via OpenMedConfig

To force MLX usage and fail fast if it is unavailable, instantiate `OpenMedConfig` from [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py) with the `backend` parameter set to `"mlx"`.

```python
from openmed import OpenMedConfig, analyze_text

cfg = OpenMedConfig(backend="mlx")   # Forces MLX; raises if unavailable

result = analyze_text(
    "Patient John Doe, DOB 1990-05-15, SSN 123-45-6789",
    model_name="pii_superclinical_large",
    config=cfg,
)
print(result.entities)

```

### Low-Level Pipeline Access

For advanced scenarios requiring direct pipeline control, import `MLXTokenClassificationPipeline` from [`openmed/mlx/inference.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/inference.py). This class mirrors the HuggingFace `pipeline("token-classification")` API, producing identical output schemas including `entity_group`, `score`, `word`, `start`, and `end` fields.

```python
from openmed.mlx.inference import MLXTokenClassificationPipeline

pipeline = MLXTokenClassificationPipeline(
    model_path="/path/to/OpenMed-PII-Model-mlx",   # Local MLX artifact

    aggregation_strategy="simple",
)
entities = pipeline("Patient Jane Roe, DOB 12/12/1980")
print(entities)

```

## MLX Artifact Contract and Cross-Platform Compatibility

The **MLX artifact contract** is documented in [`docs/mlx-backend.md`](https://github.com/maziyarpanahi/openmed/blob/main/docs/mlx-backend.md), specifying the required model directory layout, [`id2label.json`](https://github.com/maziyarpanahi/openmed/blob/main/id2label.json) mapping, and tokenizer references. When you request an MLX-only model name (e.g., `OpenMed/privacy-filter-mlx`) on a non-Apple host, the `resolve_privacy_filter_model()` function in [`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py) transparently falls back to the equivalent PyTorch checkpoint and emits a one-time `UserWarning`.

This guarantees that downstream components—including entity-merging logic, quality-gate checks, and the REST service—continue to function regardless of the underlying runtime. The fallback mechanism ensures thatMLX artifacts are portable across development environments while maintaining optimal performance on Apple Silicon.

## Summary

- Install the MLX extra using `pip install "openmed[mlx]"` to enable Apple Silicon acceleration.
- The `get_backend()` function in [`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py) auto-detects MLX availability based on platform and import checks.
- Use `OpenMedConfig(backend="mlx")` to force MLX usage and raise an error if the framework is unavailable.
- `MLXTokenClassificationPipeline` in [`openmed/mlx/inference.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/inference.py) provides a drop-in replacement for HuggingFace pipelines with identical output schemas.
- Non-Apple hosts automatically fall back to PyTorch checkpoints via `resolve_privacy_filter_model()` with a warning notification.

## Frequently Asked Questions

### Do I need to modify my code to use MLX on Apple Silicon?

No. OpenMed's `get_backend()` function automatically detects MLX availability by checking for the Darwin platform and a successful `import mlx.core`. Your existing `analyze_text()` calls will automatically route through the accelerated MLX pipeline without source code changes.

### What happens if I request an MLX-only model on a non-Apple device?

The library transparently falls back to the equivalent PyTorch checkpoint. When you specify a model like `OpenMed/privacy-filter-mlx` on a non-Darwin platform, the `resolve_privacy_filter_model()` function in [`openmed/core/backends.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/backends.py) routes to the CPU/GPU-compatible version and emits a one-time `UserWarning` about the substitution.

### How do I verify that MLX is actually being used for inference?

Confirm the dependency is installed by running `python -c "import mlx.core"`. To inspect the active backend programmatically, call `get_backend()` without arguments and check the returned class name; on Apple Silicon with the extra installed, this returns an `MLXBackend` instance.

### Can I use local MLX model artifacts instead of HuggingFace Hub downloads?

Yes. Instantiate `MLXTokenClassificationPipeline` from [`openmed/mlx/inference.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/inference.py) with a local `model_path` pointing to your MLX artifact directory containing the compiled weights and [`id2label.json`](https://github.com/maziyarpanahi/openmed/blob/main/id2label.json). This pipeline class accepts the same parameters as HuggingFace pipelines, such as `aggregation_strategy`, while loading weights directly from your filesystem.