# How to Use Zero-Shot NER with GLiNER Family Models in OpenMed

> Learn zero-shot NER with GLiNER models in OpenMed. Use the unified inference API to easily extract entities without specific training.  Install dependencies, create a request, and infer.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**OpenMed provides a unified inference API that lets you run zero-shot named entity recognition using GLiNER family models by installing optional dependencies, creating a `NerRequest`, and calling the `infer()` function which automatically routes to the appropriate model family handler.**

OpenMed ships with a dedicated zero-shot NER layer that abstracts the complexity of GLiNER and GLiNER-2 models behind a simple interface. This guide explains how to perform zero-shot NER with GLiNER family models using the library's unified `infer` API and model indexing system according to the `maziyarpanahi/openmed` source code.

## Installing GLiNER Dependencies

Before running inference, you must install the optional GLiNER dependencies, which are not included in OpenMed's core package. The helper function `is_gliner_available()` in [`openmed/ner/families/gliner.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/families/gliner.py) (lines 19-27) validates the installation and raises a `MissingDependencyError` if the packages are absent.

Install the dependencies using the extras flag:

```bash
pip install .[gliner]

```

This command installs `gliner`, `torch`, and `transformers` in a single step.

## Loading GLiNER Models with GLiNERHandle

OpenMed abstracts raw GLiNER model loading behind a lightweight `GLiNERHandle` class. The factory function `load_gliner_handle()` in [`openmed/ner/families/gliner.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/families/gliner.py) (lines 99-106) wraps `gliner.GLiNER.from_pretrained` and handles caching, authentication, and device placement.

Load a model by specifying any HuggingFace model ID from the GLiNER family:

```python
from openmed.ner.families.gliner import load_gliner_handle

handle = load_gliner_handle(
    model_id="gliner-biomed-tiny",
    cache_dir="~/.cache/huggingface",
    token="<hf-token>",  # optional, for private repositories

    device="cuda"        # or "cpu"

)

```

The function returns a handle that exposes the `predict_entities()` method for raw inference.

## Running Zero-Shot NER Inference

The high-level entry point for all NER operations is the `infer()` function in [`openmed/ner/infer.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/infer.py) (lines 70-99). This function constructs a `NerRequest`, resolves the appropriate label set, selects the family-specific runner, and returns a structured `NerResponse`.

For GLiNER models, the execution path flows through `_run_gliner_inference` which performs four critical steps:

1. **Validation**: Calls `ensure_gliner_available()` (lines 61-66) to fail fast if GLiNER is not importable.
2. **Handle Loading**: Retrieves a cached `GLiNERHandle` via `load_gliner_handle`.
3. **Prediction**: Executes `handle.predict_entities()` with the provided text, label list, and threshold (lines 74-80).
4. **Conversion**: Maps raw GLiNER output to the unified `Entity` dataclass using `_convert_gliner_entity` (lines 14-25).

### Basic Usage Example

Create a `NerRequest` and call `infer()` to extract entities without training:

```python
from openmed.ner.infer import NerRequest, infer
from openmed.core.config import OpenMedConfig

# Optional: configure HF token, cache location, and device

cfg = OpenMedConfig(
    hf_token="hf_XXXXXXXXXXXXXXXX",
    cache_dir="~/.cache/huggingface",
    device="cpu"
)

request = NerRequest(
    model_id="gliner-biomed-tiny",
    text="Patient was prescribed 5 mg of Lisinopril daily.",
    threshold=0.4,
    labels=None,  # defaults to the model's domain-specific labels

)

response = infer(request, config=cfg)

for ent in response.entities:
    print(f"- {ent.text!r} [{ent.label}] (score={ent.score:.2f})")

```

Under the hood, `infer` queries the model index in [`openmed/ner/indexing.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/indexing.py) (lines 170-176), determines that `"gliner-biomed-tiny"` belongs to the `gliner` family, and routes to `_run_gliner_inference`.

### Using GLiNER-2 Models

OpenMed supports the faster GLiNER-2 variant using the same API. The model index automatically routes requests to `_run_gliner2_inference` based on the model ID:

```python
request = NerRequest(
    model_id="gliner-uni-encoder-span",
    text="The MRI showed a hyperintense lesion in the left temporal lobe.",
    threshold=0.5,
)

response = infer(request)  # uses default config

```

The index maps `"gliner-uni-encoder-span"` to `ModelFamily.GLINER2`, ensuring the correct handler is invoked without code changes.

## Understanding the Model Index and Family Routing

OpenMed uses a centralized indexing system defined in [`openmed/ner/indexing.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/indexing.py) to map model IDs to their respective families. When you pass a `model_id` to `infer()`, the system looks up the record and instantiates the correct family-specific runner.

This architecture means switching between GLiNER and GLiNER-2 requires only changing the `model_id` in your `NerRequest`. All families share the same `NerRequest` and `NerResponse` schema, ensuring consistent behavior across model types.

## Summary

- **Install GLiNER dependencies** using `pip install .[gliner]` and verify with `is_gliner_available()`.
- **Load models** via `load_gliner_handle()` which wraps `gliner.GLiNER.from_pretrained` and handles device placement.
- **Run inference** through the unified `infer()` function in [`openmed/ner/infer.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/infer.py), which accepts a `NerRequest` and returns a `NerResponse`.
- **Leverage automatic routing** via the model index in [`openmed/ner/indexing.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/indexing.py) to switch between GLiNER and GLiNER-2 without changing your application code.
- **Adjust sensitivity** using the `threshold` parameter in `NerRequest` to control prediction confidence.

## Frequently Asked Questions

### What is the difference between GLiNER and GLiNER-2 in OpenMed?

GLiNER-2 represents a faster, optimized variant of the original GLiNER architecture. In OpenMed, both models use the same `NerRequest` interface, but the library routes GLiNER-2 models to `_run_gliner2_inference` instead of `_run_gliner_inference` based on the model index entry. The [`openmed/ner/families/gliner2.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/ner/families/gliner2.py) file contains the family-specific implementation for the newer variant.

### How do I check if GLiNER dependencies are installed correctly?

Call `openmed.ner.families.gliner.is_gliner_available()` before attempting inference. This function checks for the presence of required packages and raises a descriptive `MissingDependencyError` if they are missing, allowing you to handle missing dependencies gracefully in production code.

### Can I use custom labels with GLiNER models in OpenMed?

Yes. Pass a list of label strings to the `labels` parameter in `NerRequest`. When `labels` is `None`, OpenMed defaults to domain-specific labels defined in the model index. Custom labels enable zero-shot extraction of entity types not originally included in the model's training data.

### Where does OpenMed store downloaded GLiNER models?

OpenMed uses the HuggingFace cache directory specified in `OpenMedConfig.cache_dir`, defaulting to `~/.cache/huggingface`. The `load_gliner_handle()` function passes this path to `gliner.GLiNER.from_pretrained`, ensuring models are stored persistently and reused across sessions.