How to Implement Batch PII Extraction with BatchProcessor in OpenMed
OpenMed's BatchProcessor class enables efficient, chunk-wise PII extraction from large text collections by automatically handling model caching, smart entity merging, and error aggregation when configured with the "extract_pii" operation.
Batch PII extraction is essential for processing clinical notes and patient records at scale without exhausting memory or reloading models unnecessarily. The OpenMed library accelerates this workflow through a dedicated BatchProcessor that orchestrates the entire pipeline from input validation to result aggregation. This guide demonstrates how to implement batch PII extraction with BatchProcessor using the actual source implementation from the maziyarpanahi/openmed repository.
Understanding the Batch PII Architecture
The Core Extraction Pipeline
When the BatchProcessor operation is set to "extract_pii", it delegates heavy lifting to the private helper _extract_pii_batch defined in openmed/core/pii.py (lines 30-34). According to the OpenMed source code, this helper executes five critical steps:
- Input normalization – Optionally strips accents and resolves the language-specific model via
_resolve_effective_pii_model(lines 51-64). - Pipeline reuse – Maintains a single cached privacy-filter pipeline instance so models load only once per batch rather than per item.
- Backend execution – Calls
create_privacy_filter_pipeline(for privacy-filter models) or the standardanalyze_textfunction on the entire text list. - Smart merging – Applies
_apply_pii_smart_merging(lines 84-90) to combine overlapping entity spans into coherent PII entries. - Validation – Runs
validate_entity_spans(lines 126-130) to ensure extracted boundaries align with actual text content before returning results.
Chunking and Orchestration
The BatchProcessor class in openmed/processing/batch.py manages the high-level workflow:
- Initialization – Validates
batch_sizeusingvalidate_batch_sizefromopenmed/utils/validation.py(lines 158-174). - Chunking – Splits input lists into slices via
_iter_chunks(lines 99-103). - Per-chunk processing – Routes PII-specific logic through
_process_pii_chunk(lines 98-102). - Result aggregation – Constructs a
BatchResultobject (lines 70-106) containing per-item results, execution timings, and human-readable summaries. - Error handling – Implements
continue_on_errorlogic via_process_single_item(lines 68-72) to either halt or log failures.
Implementation Examples
Basic Batch Processing with process_batch
For rapid implementation, use the process_batch convenience function, which instantiates a BatchProcessor behind the scenes and invokes processor.process_texts() (lines 146-155).
from openmed import process_batch
texts = [
"Patient John Doe lives at 123 Main St, Springfield.",
"Contact: jane.doe@example.com, phone 555‑1234."
]
# Extract PII with default batch_size of 8 and confidence threshold of 0.5
result = process_batch(texts, operation="extract_pii")
print(result.summary())
for item in result.get_successful_results():
print(f"{item.id}: {item.result.entities}")
Customizing Batch Size and Confidence Thresholds
For production workloads requiring strict PII detection or specific throughput targets, instantiate BatchProcessor directly with custom parameters.
from openmed.processing.batch import BatchProcessor
def progress(current: int, total: int, item_result):
print(f"[{current}/{total}] {item_result.id} – "
f"{'OK' if item_result.success else 'FAIL'}")
processor = BatchProcessor(
model_name="pii_en_small",
operation="extract_pii",
batch_size=4, # Process 4 texts per chunk
confidence_threshold=0.7, # Stricter confidence cutoff
continue_on_error=False, # Stop on first failure
)
texts = ["Patient record 1...", "Patient record 2..."] # Your clinical notes
result = processor.process_texts(texts, progress_callback=progress)
The progress_callback receives the cumulative count, total items, and the BatchItemResult for real-time monitoring.
Processing Files from Directories
To batch-process entire directories of clinical notes, use process_directory, which globs files, reads UTF-8 content, and streams results through the same PII pipeline (lines 78-86).
from openmed import BatchProcessor
processor = BatchProcessor(operation="extract_pii", batch_size=5)
batch_result = processor.process_directory(
directory="data/clinical_notes",
pattern="*.txt",
recursive=True,
progress_callback=lambda cur, tot, r: print(f"{cur}/{tot} – {r.id}")
)
print(f"Successfully extracted PII from {batch_result.successful_items} files.")
Streaming Large Datasets with Iterators
For memory-constrained environments processing millions of records, use iter_process (lines 124-131) to yield results one at a time without materializing the entire batch in memory.
from openmed.processing.batch import BatchProcessor
processor = BatchProcessor(operation="extract_pii", batch_size=10)
for item_result in processor.iter_process(very_large_text_sequence):
if item_result.success:
# Process item_result.result (PredictionResult)
entities = item_result.result.entities
...
else:
# Handle or log the specific failure
print(f"Failed to process {item_result.id}: {item_result.error}")
Summary
Batch PII extraction in OpenMed leverages a sophisticated pipeline that maximizes throughput while ensuring accuracy:
- Efficient chunking via
_iter_chunksprocesses data in configurable batches to balance memory usage and speed. - Model reuse ensures the privacy-filter pipeline loads only once per batch, eliminating redundant initialization overhead.
- Smart merging combines overlapping entities automatically using
_apply_pii_smart_mergingbefore final validation. - Flexible entry points include the high-level
process_batchfunction and the configurableBatchProcessorclass for directory and iterator-based workflows. - Robust error handling via
continue_on_errorallows pipelines to complete partially even when individual texts fail processing.
Frequently Asked Questions
What is the default batch size for BatchProcessor?
The default batch_size is 8 texts per chunk. You can override this in the BatchProcessor constructor or via the process_batch helper. The system validates this value using validate_batch_size in openmed/utils/validation.py (lines 158-174) to prevent memory issues.
How does BatchProcessor handle PII model loading?
The processor reuses a single cached privacy-filter pipeline across the entire batch, as implemented in _extract_pii_batch within openmed/core/pii.py. This prevents the expensive operation of reloading transformer models for every text chunk, significantly improving throughput for large datasets.
Can I process files recursively from subdirectories?
Yes. The process_directory method accepts a recursive=True parameter that globs all matching files in subdirectories. It reads each file as UTF-8 text and processes them through the standard batch PII pipeline, returning aggregated results in a BatchResult object.
What happens if one text fails during batch processing?
Behavior depends on the continue_on_error parameter. When set to True (default), the processor logs the failure via _process_single_item and continues with remaining texts, including partial successes in the final BatchResult. When False, the entire batch operation halts immediately upon the first failure.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →