Implementing PII De-identification with Hash Redaction in OpenMed
OpenMed implements PII de-identification using SHA-256 hash redaction in openmed/core/pii.py, generating consistent 8-character hashes stored on PIIEntity objects to enable HIPAA-compliant linkability without exposing raw identifiers.
The maziyarpanahi/openmed repository provides a clinical NLP pipeline designed to protect sensitive patient information. When implementing PII de-identification with hash redaction, the library transforms identifiable entities into deterministic, cryptographic tokens that preserve relationship mapping across documents while ensuring the original text never appears in output.
How Hash Redaction Works in OpenMed
The implementation resides in openmed/core/pii.py, where the de-identification pipeline operates through three distinct stages.
The Three-Stage Pipeline
-
Extraction – The
extract_piifunction runs a token-classification model and returns aPredictionResultcontainingEntityPredictionobjects for each detected PII entity. -
Method Resolution –
_resolve_deidentification_method(lines 602-618) validates the user-requested method and converts the string"hash"into the internalDeidentificationMethodenum. -
Entity Redaction –
_redact_entity(lines 949-954) performs the actual transformation. When the method equals"hash", the function executes:
elif method == "hash":
# Generate consistent hash
hash_val = hashlib.sha256(entity.text.encode()).hexdigest()[:8]
entity.hash_value = hash_val
return f"{entity.entity_type}_{hash_val}"
Consistent Hashing and Linkability
The SHA-256 algorithm generates a consistent hash for identical input text, truncated to eight characters for readability. OpenMed stores this value in the hash_value attribute of the PIIEntity class (lines 71-78), enabling downstream systems to correlate redacted values across multiple documents without exposing the original PII.
The output follows the pattern {ENTITY_TYPE}_{HASH}, such as NAME_1a2b3c4d. This design ensures that two occurrences of the same name receive identical hashes within a single run, maintaining relationship integrity while achieving HIPAA-compliant de-identification.
Code Examples for Hash Redaction
Use the high-level deidentify function (lines 998-1008) to process text with hash redaction:
from openmed import deidentify
# Simple hash redaction
txt = "Patient John Doe visited on 01/15/2020"
res = deidentify(txt, method="hash")
print(res.deidentified_text)
# → "Patient NAME_3f9a2c1d visited on 01/15/2020"
# Access stored hash for correlation
entity = res.pii_entities[0]
print(entity.label, entity.hash_value)
# → NAME 3f9a2c1d
For debugging or custom implementations, access the low-level helper directly:
from openmed.core.pii import PIIEntity, _redact_entity
e = PIIEntity(text="John Doe", label="NAME", start=8, end=16, confidence=0.99)
redacted = _redact_entity(e, method="hash")
print(redacted) # → "NAME_a1b2c3d4"
print(e.hash_value) # → "a1b2c3d4"
Source File Architecture
The following components comprise the hash redaction system:
-
DeidentificationMethod(lines 55-56) – Typed literal defining allowed methods:"mask","remove","replace","hash", and"shift_dates". -
_resolve_deidentification_method(lines 602-618) – Normalizes user input and validates method combinations, including theshift_datesflag. -
_redact_entity(lines 949-954) – Switches on the de-identification method and implements the SHA-256 hashing logic. -
PIIEntity(lines 71-78) – ExtendsEntityPredictionwith thehash_valuefield to persist generated hashes for cross-document linking. -
openmed/cli/main.py– Exposes the--method hashoption for command-line usage. -
tests/unit/test_pii.py(lines 380-399, 589-597) – Contains unit tests verifying deterministic hash generation and properhash_valuestorage.
Summary
- OpenMed implements PII de-identification with hash redaction in
openmed/core/pii.pyusing SHA-256 cryptography. - The
deidentifyfunction acceptsmethod="hash"to trigger consistent hashing of detected entities. - Hash values are truncated to 8 characters and stored in
entity.hash_valuefor HIPAA-compliant linkability. - The output format
{ENTITY_TYPE}_{HASH}preserves entity categories while obscuring identifiers. - The pipeline guarantees that identical inputs receive identical hashes within a single execution, enabling relationship tracing without exposing raw PII.
Frequently Asked Questions
How does OpenMed ensure hash consistency across documents?
OpenMed uses SHA-256 hashing on the raw entity text (entity.text) within the _redact_entity function. Because the hash is deterministic, identical strings produce identical 8-character hex values, allowing correlation of the same patient or identifier across multiple documents without storing mapping tables.
What is the output format for hash-redacted entities?
The redacted token follows the pattern {ENTITY_TYPE}_{HASH}, such as NAME_1a2b3c4d or DATE_9f8e7d6c. This format maintains the semantic category of the PII while replacing the sensitive value with a non-reversible, consistent identifier.
Where is the hash value stored for downstream correlation?
The generated hash is stored in the hash_value attribute of the PIIEntity object (defined in openmed/core/pii.py lines 71-78). This field persists alongside the de-identified text, enabling downstream analytics to group related entities or reconstruct relationship graphs without accessing the original sensitive data.
How do I run tests to verify hash redaction behavior?
Execute the specific test suite using pytest:
pytest tests/unit/test_pii.py::test_deidentify_hash_method
The tests verify that hash redaction produces non-empty placeholders, stores values in hash_value, and generates deterministic results across multiple calls with the same input.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →