Selecting the Correct Device for OpenMed Inference: A Complete Guide

OpenMed determines which hardware accelerator to use by reading the device attribute from OpenMedConfig, supporting CPU, CUDA, MLX, and CoreML backends with automatic CPU fallback when unset.

OpenMed is an open-source medical NLP framework that abstracts hardware acceleration behind a unified configuration layer. Whether you are deploying on NVIDIA GPUs, Apple Silicon, or CoreML-compatible iOS devices, selecting the correct device for OpenMed inference requires understanding how the configuration object propagates device selection through the inference pipeline.

Where Device Configuration Is Stored

The OpenMedConfig class defines the device attribute in openmed/core/config.py (lines 59-60). This optional field accepts string values such as "cpu", "cuda", "mlx", or "coreml".

When the attribute is left as None, OpenMed defers device selection to the underlying backend libraries. This design allows the same code to run across different hardware environments without modification.

How Device Resolution Works During Inference

The inference orchestration code in openmed/ner/infer.py resolves the effective configuration before loading any model. At lines 165-166, the code retrieves the active config instance:

effective_config = config or get_config()
device = effective_config.device

This value is immediately forwarded to backend-specific loader functions. For GLiNER-based models, the device parameter passes directly to load_gliner_handle or load_gliner2_handle (lines 167-172), ensuring the model tensors move to the requested accelerator before inference begins.

Backend-Specific Hardware Support

OpenMed supports multiple execution backends, each interpreting the device string according to their respective runtimes.

PyTorch and Hugging Face

When backend="hf" (the default), the ModelLoader creates a Hugging Face pipeline that internally respects the device string. The loader receives the config's device value and passes it to the pipeline constructor, enabling CUDA acceleration with device="cuda" or CPU execution with device="cpu".

Apple MLX

For Apple Silicon deployments, setting backend="mlx" and device="mlx" routes inference through the MLX runtime. The load_gliner_handle and load_gliner2_handle functions accept the device= argument and transfer model weights to the Apple Neural Engine or unified memory.

CoreML

On-device iOS and macOS deployments use device="coreml". The inference code treats this identifier identically to other backends, though the underlying runtime utilizes the CoreML framework for optimized mobile execution.

Fallback Behavior When Device Is Unset

If OpenMedConfig.device remains None, OpenMed mirrors PyTorch's default behavior. The evaluation utilities in openmed/eval/metrics.py (line 226) explicitly default to "cpu" when no device is specified, ensuring safe execution on machines without GPU acceleration.

This fallback guarantees that inference requests succeed regardless of hardware availability, though performance will reflect CPU-based execution speeds.

Three Methods to Configure the Device

You can influence device selection through three distinct mechanisms, listed from most explicit to most implicit.

Programmatic Configuration

Create an explicit OpenMedConfig instance and pass it to the infer function. This method provides the highest granularity, allowing different devices for different requests in the same process.

from openmed import NerRequest, infer, OpenMedConfig

# Explicitly request CUDA GPU

cfg = OpenMedConfig(device="cuda")
request = NerRequest(
    model_id="disease_detection_superclinical",
    text="Patient presents with chronic myeloid leukemia.",
)
response = infer(request, config=cfg)
print(response.entities)

Environment Profile

Set the OPENMED_PROFILE environment variable to reference a TOML configuration file stored in ~/.config/openmed/profiles/. This approach centralizes hardware settings across multiple scripts or services.

import os
os.environ["OPENMED_PROFILE"] = "gpu"  # References ~/.config/openmed/profiles/gpu.toml

from openmed import analyze_text

entities = analyze_text(
    "The MRI showed a 2 cm lesion in the frontal lobe.",
    model_name="anatomy_detection_electramed",
)
print(entities)

Profile file contents (~/.config/openmed/profiles/gpu.toml):

device = "cuda"
backend = "hf"

Explicit Parameter in High-Level APIs

Some convenience functions like analyze_text accept a device parameter that forwards to the underlying config. When omitted, these functions rely on auto-detection.

from openmed import analyze_text

# Auto-detection: device=None internally, falls back to CPU

entities = analyze_text(
    "Patient received 75 mg clopidogrel for NSTEMI.",
    model_name="pharma_detection_superclinical",
)
print(entities)

Summary

  • Configuration Source: The device attribute in OpenMedConfig (openmed/core/config.py) controls hardware selection.
  • Resolution Path: The inference engine reads effective_config.device (openmed/ner/infer.py) and passes it to backend loaders.
  • Supported Backends: PyTorch/HF ("cuda", "cpu"), Apple MLX ("mlx"), and CoreML ("coreml") interpret the device string according to their respective runtimes.
  • Fallback Strategy: Unset device values default to "cpu" (openmed/eval/metrics.py), ensuring cross-platform compatibility.
  • Configuration Methods: Set devices programmatically via OpenMedConfig, through environment profiles using OPENMED_PROFILE, or via explicit parameters in high-level APIs.

Frequently Asked Questions

What happens if I don't specify a device in OpenMed?

If the device attribute is None or omitted, OpenMed defaults to CPU execution. According to the source code in openmed/eval/metrics.py (line 226), the framework explicitly assumes "cpu" when no device is provided, ensuring inference succeeds on any hardware but utilizing CPU-only computation.

Can I run OpenMed on Apple Silicon without installing CUDA?

Yes. Set OpenMedConfig(backend="mlx", device="mlx") to leverage the MLX runtime optimized for Apple Silicon. The load_gliner_handle and load_gliner2_handle functions in the inference pipeline accept the MLX device identifier and manage model placement on the Apple Neural Engine or unified memory, requiring no NVIDIA dependencies.

How do I switch between CPU and GPU for different inference requests?

Instantiate separate OpenMedConfig objects with different device values and pass them explicitly to infer(). Because the configuration is scoped per request via the effective_config = config or get_config() pattern in openmed/ner/infer.py, you can route one request to "cuda" and another to "cpu" within the same Python process without reloading the application.

Is there a performance penalty for using auto-detection?

Auto-detection itself incurs minimal overhead, but the fallback to "cpu" when no GPU is available or specified results in significantly slower inference compared to CUDA or MLX acceleration. For production deployments, explicitly set the device in OpenMedConfig or via an environment profile to avoid inadvertently running on CPU when GPU acceleration is available.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →