# Selecting the Correct Device for OpenMed Inference: A Complete Guide

> Learn how to select the correct device for OpenMed inference. Explore CPU, CUDA, MLX, and CoreML support for optimal performance with automatic fallback.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-12

---

**OpenMed determines which hardware accelerator to use by reading the `device` attribute from `OpenMedConfig`, supporting CPU, CUDA, MLX, and CoreML backends with automatic CPU fallback when unset.**

OpenMed is an open-source medical NLP framework that abstracts hardware acceleration behind a unified configuration layer. Whether you are deploying on NVIDIA GPUs, Apple Silicon, or CoreML-compatible iOS devices, selecting the correct device for OpenMed inference requires understanding how the configuration object propagates device selection through the inference pipeline.

## Where Device Configuration Is Stored

The `OpenMedConfig` class defines the `device` attribute in [`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py) (lines 59-60). This optional field accepts string values such as `"cpu"`, `"cuda"`, `"mlx"`, or `"coreml"`.

When the attribute is left as `None`, OpenMed defers device selection to the underlying backend libraries. This design allows the same code to run across different hardware environments without modification.

## How Device Resolution Works During Inference

The inference orchestration code in [`openmed/ner/infer.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/infer.py) resolves the effective configuration before loading any model. At lines 165-166, the code retrieves the active config instance:

```python
effective_config = config or get_config()
device = effective_config.device

```

This value is immediately forwarded to backend-specific loader functions. For GLiNER-based models, the device parameter passes directly to `load_gliner_handle` or `load_gliner2_handle` (lines 167-172), ensuring the model tensors move to the requested accelerator before inference begins.

## Backend-Specific Hardware Support

OpenMed supports multiple execution backends, each interpreting the `device` string according to their respective runtimes.

### PyTorch and Hugging Face

When `backend="hf"` (the default), the `ModelLoader` creates a Hugging Face pipeline that internally respects the device string. The loader receives the config's `device` value and passes it to the pipeline constructor, enabling CUDA acceleration with `device="cuda"` or CPU execution with `device="cpu"`.

### Apple MLX

For Apple Silicon deployments, setting `backend="mlx"` and `device="mlx"` routes inference through the MLX runtime. The `load_gliner_handle` and `load_gliner2_handle` functions accept the `device=` argument and transfer model weights to the Apple Neural Engine or unified memory.

### CoreML

On-device iOS and macOS deployments use `device="coreml"`. The inference code treats this identifier identically to other backends, though the underlying runtime utilizes the CoreML framework for optimized mobile execution.

## Fallback Behavior When Device Is Unset

If `OpenMedConfig.device` remains `None`, OpenMed mirrors PyTorch's default behavior. The evaluation utilities in [`openmed/eval/metrics.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/eval/metrics.py) (line 226) explicitly default to `"cpu"` when no device is specified, ensuring safe execution on machines without GPU acceleration.

This fallback guarantees that inference requests succeed regardless of hardware availability, though performance will reflect CPU-based execution speeds.

## Three Methods to Configure the Device

You can influence device selection through three distinct mechanisms, listed from most explicit to most implicit.

### Programmatic Configuration

Create an explicit `OpenMedConfig` instance and pass it to the `infer` function. This method provides the highest granularity, allowing different devices for different requests in the same process.

```python
from openmed import NerRequest, infer, OpenMedConfig

# Explicitly request CUDA GPU

cfg = OpenMedConfig(device="cuda")
request = NerRequest(
    model_id="disease_detection_superclinical",
    text="Patient presents with chronic myeloid leukemia.",
)
response = infer(request, config=cfg)
print(response.entities)

```

### Environment Profile

Set the `OPENMED_PROFILE` environment variable to reference a TOML configuration file stored in `~/.config/openmed/profiles/`. This approach centralizes hardware settings across multiple scripts or services.

```python
import os
os.environ["OPENMED_PROFILE"] = "gpu"  # References ~/.config/openmed/profiles/gpu.toml

from openmed import analyze_text

entities = analyze_text(
    "The MRI showed a 2 cm lesion in the frontal lobe.",
    model_name="anatomy_detection_electramed",
)
print(entities)

```

Profile file contents (`~/.config/openmed/profiles/gpu.toml`):

```toml
device = "cuda"
backend = "hf"

```

### Explicit Parameter in High-Level APIs

Some convenience functions like `analyze_text` accept a `device` parameter that forwards to the underlying config. When omitted, these functions rely on auto-detection.

```python
from openmed import analyze_text

# Auto-detection: device=None internally, falls back to CPU

entities = analyze_text(
    "Patient received 75 mg clopidogrel for NSTEMI.",
    model_name="pharma_detection_superclinical",
)
print(entities)

```

## Summary

- **Configuration Source**: The `device` attribute in `OpenMedConfig` ([`openmed/core/config.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/config.py)) controls hardware selection.
- **Resolution Path**: The inference engine reads `effective_config.device` ([`openmed/ner/infer.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/infer.py)) and passes it to backend loaders.
- **Supported Backends**: PyTorch/HF (`"cuda"`, `"cpu"`), Apple MLX (`"mlx"`), and CoreML (`"coreml"`) interpret the device string according to their respective runtimes.
- **Fallback Strategy**: Unset device values default to `"cpu"` ([`openmed/eval/metrics.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/eval/metrics.py)), ensuring cross-platform compatibility.
- **Configuration Methods**: Set devices programmatically via `OpenMedConfig`, through environment profiles using `OPENMED_PROFILE`, or via explicit parameters in high-level APIs.

## Frequently Asked Questions

### What happens if I don't specify a device in OpenMed?

If the `device` attribute is `None` or omitted, OpenMed defaults to CPU execution. According to the source code in [`openmed/eval/metrics.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/eval/metrics.py) (line 226), the framework explicitly assumes `"cpu"` when no device is provided, ensuring inference succeeds on any hardware but utilizing CPU-only computation.

### Can I run OpenMed on Apple Silicon without installing CUDA?

Yes. Set `OpenMedConfig(backend="mlx", device="mlx")` to leverage the MLX runtime optimized for Apple Silicon. The `load_gliner_handle` and `load_gliner2_handle` functions in the inference pipeline accept the MLX device identifier and manage model placement on the Apple Neural Engine or unified memory, requiring no NVIDIA dependencies.

### How do I switch between CPU and GPU for different inference requests?

Instantiate separate `OpenMedConfig` objects with different `device` values and pass them explicitly to `infer()`. Because the configuration is scoped per request via the `effective_config = config or get_config()` pattern in [`openmed/ner/infer.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/infer.py), you can route one request to `"cuda"` and another to `"cpu"` within the same Python process without reloading the application.

### Is there a performance penalty for using auto-detection?

Auto-detection itself incurs minimal overhead, but the fallback to `"cpu"` when no GPU is available or specified results in significantly slower inference compared to CUDA or MLX acceleration. For production deployments, explicitly set the device in `OpenMedConfig` or via an environment profile to avoid inadvertently running on CPU when GPU acceleration is available.