Ensuring HIPAA Compliance with On-Device PII Detection in OpenMed
OpenMed ensures HIPAA compliance by running all PII detection, merging, and de-identification locally on the device, utilizing extract_pii and deidentify functions from openmed/core/pii.py with support for all 18 Safe Harbor identifiers.
The OpenMed library provides a complete on-device solution for detecting and removing personally identifiable information (PII) from clinical text. By processing all protected health information (PHI) locally without network transmission, the library satisfies strict HIPAA Safe Harbor requirements while maintaining full auditability. This architecture ensures that sensitive patient data never leaves the device during the de-identification pipeline.
On-Device PII Detection Architecture
PII Extraction in openmed/core/pii.py
The extraction pipeline begins with the extract_pii helper function, exposed as a top-level API import. This function leverages existing NER infrastructure for model loading, tokenization, and inference, then applies language-specific post-processing targeting the 18 HIPAA Safe Harbor identifiers.
Detected entities are represented by the PIIEntity dataclass defined in lines 59-77 of openmed/core/pii.py. Each entity captures the text span, label type, confidence score, and positional indices required for downstream processing.
Smart Entity Merging
Raw token-level predictions often fragment multi-token entities. OpenMed consolidates these spans using the entity-merger module located in openmed/core/pii_entity_merger.py.
The PII_PATTERNS dictionary (lines 42-150) defines regex-based semantic units for dates, SSNs, phone numbers, and other identifiers. Each pattern applies a low base confidence combined with a context boost mechanism modeled after Microsoft Presidio, ensuring only medically relevant matches survive filtering.
Validation helpers such as validate_ssn and validate_luhn enforce official checksum rules, preventing false positives on invalid identifier formats.
De-Identification Strategies
The deidentify function (also exported from the package top level) converts merged entities into HIPAA-compliant redactions. Supported strategies are enumerated by the DeidentificationMethod class (lines 55-57) and include:
- mask: Replace with entity type label (e.g.,
[NAME]) - remove: Delete the text entirely
- replace: Substitute with synthetic placeholders (reversible mapping)
- hash: Cryptographic hashing
- shift_dates: Temporal shifting for date de-identification
The operation returns a DeidentificationResult object (lines 86-104) containing the redacted text, a mapping of original to redacted strings, and metadata for audit trails.
HIPAA Safe Harbor Compliance Implementation
OpenMed satisfies HIPAA requirements through the following technical safeguards:
| HIPAA Requirement | OpenMed Implementation |
|---|---|
| All 18 Safe Harbor identifiers | PII_PATTERNS covers every identifier including names, dates, SSNs, phone numbers, and geographic data |
| On-device processing | All detection, merging, and redaction execute locally; no PHI transmitted over networks |
| Reversible de-identification | The mapping field in DeidentificationResult stores reversible links when using the replace method |
| Configurable thresholds | Users can tune context-boost scores or filter entities below confidence cutoffs |
| Auditability | DeidentificationResult.timestamp and metadata provide traceable records for compliance audits |
Practical Implementation Examples
The following examples demonstrate the complete workflow using the public API.
First, extract PII entities from a clinical note:
from openmed import extract_pii, deidentify
clinical_note = """
Patient John Doe was admitted on 01/15/1970. His SSN is 123-45-6789.
Contact: (555) 123‑4567. Discharged on 02/20/1970.
"""
# Detect PII entities (returns a list of PIIEntity objects)
pii_entities = extract_pii(clinical_note).entities
for e in pii_entities:
print(f"{e.label}: {e.text} (confidence={e.confidence:.2f})")
Output:
NAME: John Doe (confidence=0.97)
DATE: 01/15/1970 (confidence=0.88)
SSN: 123-45-6789 (confidence=0.92)
PHONE: (555) 123‑4567 (confidence=0.94)
DATE: 02/20/1970 (confidence=0.86)
Next, de-identify using the mask strategy:
# De-identify the note using the "mask" strategy
result = deidentify(
clinical_note,
method="mask", # options: mask, remove, replace, hash, shift_dates
keep_year=True, # keep the year part of dates (HIPAA‑allowed)
)
print(result.deidentified_text)
Output:
Patient [NAME] was admitted on [DATE]/1970. His SSN is [SSN].
Contact: [PHONE]. Discharged on [DATE]/1970.
Finally, inspect the structured result for audit logging:
from pprint import pprint
pprint(result.to_dict())
{
"original_text": "...",
"deidentified_text": "...",
"pii_entities": [
{
"text": "John Doe",
"label": "NAME",
"entity_type": "NAME",
"start": 9,
"end": 17,
"confidence": 0.97,
"redacted_text": "[NAME]",
"metadata": {}
}
],
"method": "mask",
"timestamp": "2026-06-12T14:23:01.123456",
"num_entities_redacted": 5,
"metadata": {}
}
Summary
- On-device processing in OpenMed ensures HIPAA compliance by keeping all PHI detection and redaction local to the device.
- The
extract_piifunction inopenmed/core/pii.pyidentifies all 18 Safe Harbor identifiers using NER models and language-specific post-processing. - Smart merging via
openmed/core/pii_entity_merger.pyconsolidates fragmented spans using regex patterns and validation helpers likevalidate_ssn. deidentifysupports multiple strategies (mask, remove, replace, hash, shift_dates) and returns audit-readyDeidentificationResultobjects.- Configurable confidence thresholds and reversible mappings support both strict anonymization and research use cases requiring re-identification.
Frequently Asked Questions
How does OpenMed ensure no data leaves the device during PII detection?
All processing occurs within the local Python environment using the extract_pii and deidentify functions. The library loads NER models and regex patterns locally in openmed/core/pii.py and openmed/core/pii_entity_merger.py, performing inference and redaction without network calls or external API dependencies.
Which de-identification methods does OpenMed support?
OpenMed supports five strategies defined in the DeidentificationMethod enumeration: mask (entity type labels), remove (deletion), replace (reversible synthetic placeholders), hash (cryptographic hashing), and shift_dates (temporal offsetting). The replace method preserves mappings in DeidentificationResult.mapping for controlled re-identification.
Can OpenMed handle all 18 HIPAA Safe Harbor identifiers?
Yes. The PII_PATTERNS configuration in openmed/core/pii_entity_merger.py (lines 42-150) covers all 18 Safe Harbor categories including names, dates, SSNs, phone numbers, fax numbers, email addresses, geographic subdivisions, and medical record numbers. Validation helpers ensure patterns conform to official checksum rules.
How does the entity merging improve detection accuracy?
The pii_entity_merger.py module consolidates fragmented token-level NER predictions into semantic units using regex patterns. It applies context-boost scoring similar to Microsoft Presidio and validates identifiers using checksum algorithms (Luhn for SSNs), filtering out false positives while preserving true matches.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →