# How to Use Privacy Filter Family Models for PII Detection in OpenMed

> Easily detect PII in OpenMed using Privacy Filter models. Leverage MLX or PyTorch for token classification and get standardized JSON results with entity labels and confidence scores.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**OpenMed provides Privacy Filter family models as token-classification pipelines that detect personally identifiable information (PII) using either MLX for Apple Silicon or PyTorch for cross-platform deployment, returning standardized JSON outputs with entity labels, confidence scores, and character offsets.**

OpenMed ships the Privacy Filter family—including the original OpenAI model and multilingual variants—as a unified solution for PII detection in clinical text. These models identify sensitive spans such as names, dates, and addresses, supporting both high-performance inference on macOS via MLX and standard execution on CPU or GPU via PyTorch. The library automatically normalizes outputs across backends to a consistent schema containing entity groups, confidence scores, and character offsets.

## Architecture and Backend Options

The Privacy Filter family supports two distinct inference backends selected automatically based on your hardware and the model manifest.

### MLX Backend for Apple Silicon

The **MLX backend** delivers fast on-device inference without requiring a GPU. The core implementation resides in [`openmed/mlx/inference.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/inference.py), where the `PrivacyFilterMLXPipeline` class handles model loading and decoding.

In [`openmed/mlx/models/privacy_filter.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/models/privacy_filter.py), the model architecture includes the full transformer stack with RMSNorm, RoPE, local attention, and MoE feed-forward layers, culminating in the `OpenAIPrivacyFilterForTokenClassification` token-classification head. The pipeline uses `tiktoken` for byte-offset decoding and applies a Viterbi decoder to resolve BIOES tag sequences into coherent spans.

### Torch Backend for Cross-Platform Deployment

The **Torch backend** provides broad compatibility across operating systems and hardware. Implemented in [`openmed/torch/privacy_filter.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/torch/privacy_filter.py), the `PrivacyFilterTorchPipeline` wraps HuggingFace `AutoModelForTokenClassification` with security safeguards.

This pipeline enforces a trusted-remote-code allow-list before loading models. The constructor validates `trust_remote_code=True` against the `TRUSTED_REMOTE_CODE_MODELS` list and the `OPENMED_TRUSTED_REMOTE_CODE_MODELS` environment variable, raising `ValueError` for unauthorized repositories to prevent arbitrary code execution.

## Loading a Privacy Filter Model

Both pipelines follow an identical contract, accepting text input and returning a list of dictionaries with keys `entity_group`, `score`, `word`, `start`, and `end`.

### Using the MLX Pipeline (macOS/Apple Silicon)

For Apple Silicon devices, use the `create_mlx_pipeline` function from `openmed.mlx.inference`. This function resolves model IDs to pre-converted MLX checkpoints and initializes the Viterbi decoder.

```python
from openmed.mlx.inference import create_mlx_pipeline

# Load a multilingual privacy-filter model

pipeline = create_mlx_pipeline(
    model_name="OpenMed/privacy-filter-multilingual-mlx",
    aggregation_strategy="simple",  # Groups BIOES tokens into spans

)

text = "Patient John Doe was admitted on 2023-04-15 and lives at 123 Main St."
entities = pipeline(text)

print(entities)

```

**Output:**

```json
[
  {"entity_group": "NAME", "score": 0.99, "word": "John Doe", "start": 8, "end": 16},
  {"entity_group": "DATE", "score": 0.97, "word": "2023-04-15", "start": 32, "end": 42},
  {"entity_group": "ADDRESS", "score": 0.94, "word": "123 Main St.", "start": 53, "end": 66}
]

```

### Using the Torch Pipeline (CPU/GPU)

For cross-platform deployment or GPU acceleration, instantiate `PrivacyFilterTorchPipeline` from `openmed.torch.privacy_filter`.

```python
from openmed.torch import PrivacyFilterTorchPipeline

# Load the original OpenAI checkpoint

pipeline = PrivacyFilterTorchPipeline(
    model_name="openai/privacy-filter",
    device="cuda",  # or "cpu"

    aggregation_strategy="simple",
    trust_remote_code=False,  # Safe default

)

texts = [
    "Dr. Alice Smith prescribed medication to patient Bob.",
    "The visit was on 12/01/2022."
]

batch_entities = pipeline(texts)  # Returns list-of-lists

print(batch_entities)

```

**Output:**

```json
[
  [
    {"entity_group": "PROFESSION", "score": 0.95, "word": "Dr.", "start": 0, "end": 3},
    {"entity_group": "NAME", "score": 0.99, "word": "Alice Smith", "start": 4, "end": 15},
    {"entity_group": "NAME", "score": 0.97, "word": "Bob", "start": 46, "end": 49}
  ],
  [
    {"entity_group": "DATE", "score": 0.96, "word": "12/01/2022", "start": 16, "end": 26}
  ]
]

```

## Backend Dispatcher and Transparent Switching

The `openmed.core.backends` module abstracts backend selection through the `get_backend` function. This dispatcher checks the model manifest for `"family": "openai-privacy-filter"` and returns either an `MLXBackend` or `TorchBackend` instance.

```python
from openmed.core.backends import get_backend

# Automatically selects MLX on Apple Silicon or Torch otherwise

backend = get_backend("openai-privacy-filter")
pipeline = backend.create_pipeline(
    model_name="OpenMed/privacy-filter-multilingual",
    aggregation_strategy="simple"
)

entities = pipeline("Patient data: Jane Doe, 01-02-1990, 555-1234.")

```

Both backend classes expose an identical `create_pipeline` method, ensuring consistent behavior regardless of the underlying implementation.

## Post-Processing and Span Refinement

After initial inference, both pipelines invoke shared utilities in [`openmed/core/decoding/spans.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/decoding/spans.py) to refine raw predictions. The `trim_span_whitespace` function removes leading and trailing whitespace from detected entities, while `refine_privacy_filter_span` applies domain-specific corrections such as extending date spans or handling hyphenated names.

These post-processing steps ensure that character offsets (`start` and `end`) align precisely with the input text boundaries, maintaining consistency across the MLX and Torch codepaths.

## Summary

- **Privacy Filter family models** in OpenMed detect PII through token-classification pipelines available in both MLX and Torch variants.
- **MLX inference** on Apple Silicon uses `create_mlx_pipeline` from `openmed.mlx.inference` with Viterbi decoding via `PrivacyFilterMLXPipeline`.
- **Torch inference** uses `PrivacyFilterTorchPipeline` from `openmed.torch.privacy_filter` with strict `trust_remote_code` validation for security.
- **Backend dispatcher** in `openmed.core.backends` automatically selects the optimal implementation based on model family and hardware availability.
- **Output format** is standardized across backends, returning dictionaries with `entity_group`, `score`, `word`, `start`, and `end` keys.
- **Span refinement** occurs in [`openmed/core/decoding/spans.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/decoding/spans.py) to trim whitespace and apply label-specific corrections.

## Frequently Asked Questions

### What output format do Privacy Filter pipelines return?

Both MLX and Torch pipelines return a list of dictionaries, where each dictionary represents a detected PII span with keys including `entity_group` (the PII label), `score` (confidence probability), `word` (the detected text), and character offsets `start` and `end`. This unified schema ensures compatibility with OpenMed's de-identification utilities regardless of which backend generated the predictions.

### How do I enable trust_remote_code for custom fine-tuned models?

Set the environment variable `OPENMED_TRUSTED_REMOTE_CODE_MODELS` to your repository ID before instantiating `PrivacyFilterTorchPipeline`, then pass `trust_remote_code=True` in the constructor. The pipeline validates your repository against this allow-list before loading to prevent execution of untrusted code. Models not on the list will raise a `ValueError` if remote code execution is requested.

### Which backend should I use for production inference?

Use the **MLX backend** via `create_mlx_pipeline` when running on macOS with Apple Silicon, as it provides optimized on-device performance without requiring a GPU. For Linux servers, Windows workstations, or CUDA-accelerated inference, use the **Torch backend** via `PrivacyFilterTorchPipeline`. The `get_backend` dispatcher in `openmed.core.backends` can automatically select the appropriate implementation based on your runtime environment.

### How does OpenMed handle multilingual PII detection?

OpenMed provides multilingual variants of the Privacy Filter family, such as `OpenMed/privacy-filter-multilingual-mlx`, which use the same architecture as the original OpenAI model but support non-English clinical text. The tokenization and Viterbi decoding logic in [`openmed/mlx/inference.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/inference.py) handles byte-offset calculations for multilingual inputs, while the post-processing utilities in [`openmed/core/decoding/spans.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/decoding/spans.py) apply language-agnostic whitespace trimming and span refinement.