How to Use Privacy Filter Family Models for PII Detection in OpenMed
OpenMed provides Privacy Filter family models as token-classification pipelines that detect personally identifiable information (PII) using either MLX for Apple Silicon or PyTorch for cross-platform deployment, returning standardized JSON outputs with entity labels, confidence scores, and character offsets.
OpenMed ships the Privacy Filter family—including the original OpenAI model and multilingual variants—as a unified solution for PII detection in clinical text. These models identify sensitive spans such as names, dates, and addresses, supporting both high-performance inference on macOS via MLX and standard execution on CPU or GPU via PyTorch. The library automatically normalizes outputs across backends to a consistent schema containing entity groups, confidence scores, and character offsets.
Architecture and Backend Options
The Privacy Filter family supports two distinct inference backends selected automatically based on your hardware and the model manifest.
MLX Backend for Apple Silicon
The MLX backend delivers fast on-device inference without requiring a GPU. The core implementation resides in openmed/mlx/inference.py, where the PrivacyFilterMLXPipeline class handles model loading and decoding.
In openmed/mlx/models/privacy_filter.py, the model architecture includes the full transformer stack with RMSNorm, RoPE, local attention, and MoE feed-forward layers, culminating in the OpenAIPrivacyFilterForTokenClassification token-classification head. The pipeline uses tiktoken for byte-offset decoding and applies a Viterbi decoder to resolve BIOES tag sequences into coherent spans.
Torch Backend for Cross-Platform Deployment
The Torch backend provides broad compatibility across operating systems and hardware. Implemented in openmed/torch/privacy_filter.py, the PrivacyFilterTorchPipeline wraps HuggingFace AutoModelForTokenClassification with security safeguards.
This pipeline enforces a trusted-remote-code allow-list before loading models. The constructor validates trust_remote_code=True against the TRUSTED_REMOTE_CODE_MODELS list and the OPENMED_TRUSTED_REMOTE_CODE_MODELS environment variable, raising ValueError for unauthorized repositories to prevent arbitrary code execution.
Loading a Privacy Filter Model
Both pipelines follow an identical contract, accepting text input and returning a list of dictionaries with keys entity_group, score, word, start, and end.
Using the MLX Pipeline (macOS/Apple Silicon)
For Apple Silicon devices, use the create_mlx_pipeline function from openmed.mlx.inference. This function resolves model IDs to pre-converted MLX checkpoints and initializes the Viterbi decoder.
from openmed.mlx.inference import create_mlx_pipeline
# Load a multilingual privacy-filter model
pipeline = create_mlx_pipeline(
model_name="OpenMed/privacy-filter-multilingual-mlx",
aggregation_strategy="simple", # Groups BIOES tokens into spans
)
text = "Patient John Doe was admitted on 2023-04-15 and lives at 123 Main St."
entities = pipeline(text)
print(entities)
Output:
[
{"entity_group": "NAME", "score": 0.99, "word": "John Doe", "start": 8, "end": 16},
{"entity_group": "DATE", "score": 0.97, "word": "2023-04-15", "start": 32, "end": 42},
{"entity_group": "ADDRESS", "score": 0.94, "word": "123 Main St.", "start": 53, "end": 66}
]
Using the Torch Pipeline (CPU/GPU)
For cross-platform deployment or GPU acceleration, instantiate PrivacyFilterTorchPipeline from openmed.torch.privacy_filter.
from openmed.torch import PrivacyFilterTorchPipeline
# Load the original OpenAI checkpoint
pipeline = PrivacyFilterTorchPipeline(
model_name="openai/privacy-filter",
device="cuda", # or "cpu"
aggregation_strategy="simple",
trust_remote_code=False, # Safe default
)
texts = [
"Dr. Alice Smith prescribed medication to patient Bob.",
"The visit was on 12/01/2022."
]
batch_entities = pipeline(texts) # Returns list-of-lists
print(batch_entities)
Output:
[
[
{"entity_group": "PROFESSION", "score": 0.95, "word": "Dr.", "start": 0, "end": 3},
{"entity_group": "NAME", "score": 0.99, "word": "Alice Smith", "start": 4, "end": 15},
{"entity_group": "NAME", "score": 0.97, "word": "Bob", "start": 46, "end": 49}
],
[
{"entity_group": "DATE", "score": 0.96, "word": "12/01/2022", "start": 16, "end": 26}
]
]
Backend Dispatcher and Transparent Switching
The openmed.core.backends module abstracts backend selection through the get_backend function. This dispatcher checks the model manifest for "family": "openai-privacy-filter" and returns either an MLXBackend or TorchBackend instance.
from openmed.core.backends import get_backend
# Automatically selects MLX on Apple Silicon or Torch otherwise
backend = get_backend("openai-privacy-filter")
pipeline = backend.create_pipeline(
model_name="OpenMed/privacy-filter-multilingual",
aggregation_strategy="simple"
)
entities = pipeline("Patient data: Jane Doe, 01-02-1990, 555-1234.")
Both backend classes expose an identical create_pipeline method, ensuring consistent behavior regardless of the underlying implementation.
Post-Processing and Span Refinement
After initial inference, both pipelines invoke shared utilities in openmed/core/decoding/spans.py to refine raw predictions. The trim_span_whitespace function removes leading and trailing whitespace from detected entities, while refine_privacy_filter_span applies domain-specific corrections such as extending date spans or handling hyphenated names.
These post-processing steps ensure that character offsets (start and end) align precisely with the input text boundaries, maintaining consistency across the MLX and Torch codepaths.
Summary
- Privacy Filter family models in OpenMed detect PII through token-classification pipelines available in both MLX and Torch variants.
- MLX inference on Apple Silicon uses
create_mlx_pipelinefromopenmed.mlx.inferencewith Viterbi decoding viaPrivacyFilterMLXPipeline. - Torch inference uses
PrivacyFilterTorchPipelinefromopenmed.torch.privacy_filterwith stricttrust_remote_codevalidation for security. - Backend dispatcher in
openmed.core.backendsautomatically selects the optimal implementation based on model family and hardware availability. - Output format is standardized across backends, returning dictionaries with
entity_group,score,word,start, andendkeys. - Span refinement occurs in
openmed/core/decoding/spans.pyto trim whitespace and apply label-specific corrections.
Frequently Asked Questions
What output format do Privacy Filter pipelines return?
Both MLX and Torch pipelines return a list of dictionaries, where each dictionary represents a detected PII span with keys including entity_group (the PII label), score (confidence probability), word (the detected text), and character offsets start and end. This unified schema ensures compatibility with OpenMed's de-identification utilities regardless of which backend generated the predictions.
How do I enable trust_remote_code for custom fine-tuned models?
Set the environment variable OPENMED_TRUSTED_REMOTE_CODE_MODELS to your repository ID before instantiating PrivacyFilterTorchPipeline, then pass trust_remote_code=True in the constructor. The pipeline validates your repository against this allow-list before loading to prevent execution of untrusted code. Models not on the list will raise a ValueError if remote code execution is requested.
Which backend should I use for production inference?
Use the MLX backend via create_mlx_pipeline when running on macOS with Apple Silicon, as it provides optimized on-device performance without requiring a GPU. For Linux servers, Windows workstations, or CUDA-accelerated inference, use the Torch backend via PrivacyFilterTorchPipeline. The get_backend dispatcher in openmed.core.backends can automatically select the appropriate implementation based on your runtime environment.
How does OpenMed handle multilingual PII detection?
OpenMed provides multilingual variants of the Privacy Filter family, such as OpenMed/privacy-filter-multilingual-mlx, which use the same architecture as the original OpenAI model but support non-English clinical text. The tokenization and Viterbi decoding logic in openmed/mlx/inference.py handles byte-offset calculations for multilingual inputs, while the post-processing utilities in openmed/core/decoding/spans.py apply language-agnostic whitespace trimming and span refinement.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →