Setting up MLX Backend on Apple Silicon for Accelerated Inference with OpenMed
OpenMed automatically utilizes the Apple MLX backend on Apple Silicon devices, delivering 24-33× speed-ups over CPU-only PyTorch inference without requiring any code changes.
OpenMed provides a pluggable inference-backend abstraction that automatically selects the fastest available runtime for medical NLP workloads. When running on macOS, iPadOS, or iOS with Apple Silicon chips, the library prioritizes the Apple MLX framework for token classification and privacy filtering tasks. This guide explains how to configure and verify the MLX backend for accelerated inference using the OpenMed source code.
How OpenMed Auto-Detects the MLX Backend
The backend selection logic lives in openmed/core/backends.py. The get_backend() function queries MLXBackend.is_available() to determine whether the host supports MLX acceleration. This check returns True only when the platform is Darwin (macOS/iOS) and the mlx Python package can be imported successfully.
When these conditions are met, get_backend() automatically selects the MLX backend; otherwise, it falls back to the HuggingFace/PyTorch backend. This auto-detection happens at runtime, ensuring that scripts written on Apple Silicon will seamlessly transition to CPU-based inference on Linux or Windows hosts without modification.
from openmed.core.backends import get_backend
# Auto-detect: returns MLXBackend on Apple Silicon, PyTorchBackend elsewhere
backend = get_backend()
# Force specific backend (raises RuntimeError if MLX unavailable)
backend = get_backend(name="mlx")
Installing MLX Dependencies
MLX support is shipped as an optional extra to keep base installations lightweight. The dependency specification resides in pyproject.toml under [project.optional-dependencies], where the mlx extra pulls in mlx, transformers, tokenizers, safetensors, and huggingface-hub.
Install the accelerated backend on your Apple Silicon machine:
# From PyPI
pip install "openmed[mlx]"
# Or from source
pip install -e ".[mlx]"
Verify the installation by confirming that MLX is importable:
python -c "import mlx.core; print('MLX available')"
Using the MLX Backend
Automatic Backend Selection
Once installed, the standard OpenMed API automatically routes inference through the MLX pipeline. The analyze_text() function utilizes the auto-detected backend without additional configuration.
from openmed import analyze_text
result = analyze_text(
"Patient John Doe, DOB 1990-05-15, SSN 123-45-6789",
model_name="pii_superclinical_large", # Any MLX-compatible model
)
print(result.entities) # Uses MLX acceleration on Apple Silicon
Explicit Configuration via OpenMedConfig
To force MLX usage and fail fast if it is unavailable, instantiate OpenMedConfig from openmed/core/config.py with the backend parameter set to "mlx".
from openmed import OpenMedConfig, analyze_text
cfg = OpenMedConfig(backend="mlx") # Forces MLX; raises if unavailable
result = analyze_text(
"Patient John Doe, DOB 1990-05-15, SSN 123-45-6789",
model_name="pii_superclinical_large",
config=cfg,
)
print(result.entities)
Low-Level Pipeline Access
For advanced scenarios requiring direct pipeline control, import MLXTokenClassificationPipeline from openmed/mlx/inference.py. This class mirrors the HuggingFace pipeline("token-classification") API, producing identical output schemas including entity_group, score, word, start, and end fields.
from openmed.mlx.inference import MLXTokenClassificationPipeline
pipeline = MLXTokenClassificationPipeline(
model_path="/path/to/OpenMed-PII-Model-mlx", # Local MLX artifact
aggregation_strategy="simple",
)
entities = pipeline("Patient Jane Roe, DOB 12/12/1980")
print(entities)
MLX Artifact Contract and Cross-Platform Compatibility
The MLX artifact contract is documented in docs/mlx-backend.md, specifying the required model directory layout, id2label.json mapping, and tokenizer references. When you request an MLX-only model name (e.g., OpenMed/privacy-filter-mlx) on a non-Apple host, the resolve_privacy_filter_model() function in openmed/core/backends.py transparently falls back to the equivalent PyTorch checkpoint and emits a one-time UserWarning.
This guarantees that downstream components—including entity-merging logic, quality-gate checks, and the REST service—continue to function regardless of the underlying runtime. The fallback mechanism ensures thatMLX artifacts are portable across development environments while maintaining optimal performance on Apple Silicon.
Summary
- Install the MLX extra using
pip install "openmed[mlx]"to enable Apple Silicon acceleration. - The
get_backend()function inopenmed/core/backends.pyauto-detects MLX availability based on platform and import checks. - Use
OpenMedConfig(backend="mlx")to force MLX usage and raise an error if the framework is unavailable. MLXTokenClassificationPipelineinopenmed/mlx/inference.pyprovides a drop-in replacement for HuggingFace pipelines with identical output schemas.- Non-Apple hosts automatically fall back to PyTorch checkpoints via
resolve_privacy_filter_model()with a warning notification.
Frequently Asked Questions
Do I need to modify my code to use MLX on Apple Silicon?
No. OpenMed's get_backend() function automatically detects MLX availability by checking for the Darwin platform and a successful import mlx.core. Your existing analyze_text() calls will automatically route through the accelerated MLX pipeline without source code changes.
What happens if I request an MLX-only model on a non-Apple device?
The library transparently falls back to the equivalent PyTorch checkpoint. When you specify a model like OpenMed/privacy-filter-mlx on a non-Darwin platform, the resolve_privacy_filter_model() function in openmed/core/backends.py routes to the CPU/GPU-compatible version and emits a one-time UserWarning about the substitution.
How do I verify that MLX is actually being used for inference?
Confirm the dependency is installed by running python -c "import mlx.core". To inspect the active backend programmatically, call get_backend() without arguments and check the returned class name; on Apple Silicon with the extra installed, this returns an MLXBackend instance.
Can I use local MLX model artifacts instead of HuggingFace Hub downloads?
Yes. Instantiate MLXTokenClassificationPipeline from openmed/mlx/inference.py with a local model_path pointing to your MLX artifact directory containing the compiled weights and id2label.json. This pipeline class accepts the same parameters as HuggingFace pipelines, such as aggregation_strategy, while loading weights directly from your filesystem.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →