How to Implement Batch PII Extraction with BatchProcessor in OpenMed

OpenMed's BatchProcessor class enables efficient, chunk-wise PII extraction from large text collections by automatically handling model caching, smart entity merging, and error aggregation when configured with the "extract_pii" operation.

Batch PII extraction is essential for processing clinical notes and patient records at scale without exhausting memory or reloading models unnecessarily. The OpenMed library accelerates this workflow through a dedicated BatchProcessor that orchestrates the entire pipeline from input validation to result aggregation. This guide demonstrates how to implement batch PII extraction with BatchProcessor using the actual source implementation from the maziyarpanahi/openmed repository.

Understanding the Batch PII Architecture

The Core Extraction Pipeline

When the BatchProcessor operation is set to "extract_pii", it delegates heavy lifting to the private helper _extract_pii_batch defined in openmed/core/pii.py (lines 30-34). According to the OpenMed source code, this helper executes five critical steps:

  1. Input normalization – Optionally strips accents and resolves the language-specific model via _resolve_effective_pii_model (lines 51-64).
  2. Pipeline reuse – Maintains a single cached privacy-filter pipeline instance so models load only once per batch rather than per item.
  3. Backend execution – Calls create_privacy_filter_pipeline (for privacy-filter models) or the standard analyze_text function on the entire text list.
  4. Smart merging – Applies _apply_pii_smart_merging (lines 84-90) to combine overlapping entity spans into coherent PII entries.
  5. Validation – Runs validate_entity_spans (lines 126-130) to ensure extracted boundaries align with actual text content before returning results.

Chunking and Orchestration

The BatchProcessor class in openmed/processing/batch.py manages the high-level workflow:

  • Initialization – Validates batch_size using validate_batch_size from openmed/utils/validation.py (lines 158-174).
  • Chunking – Splits input lists into slices via _iter_chunks (lines 99-103).
  • Per-chunk processing – Routes PII-specific logic through _process_pii_chunk (lines 98-102).
  • Result aggregation – Constructs a BatchResult object (lines 70-106) containing per-item results, execution timings, and human-readable summaries.
  • Error handling – Implements continue_on_error logic via _process_single_item (lines 68-72) to either halt or log failures.

Implementation Examples

Basic Batch Processing with process_batch

For rapid implementation, use the process_batch convenience function, which instantiates a BatchProcessor behind the scenes and invokes processor.process_texts() (lines 146-155).

from openmed import process_batch

texts = [
    "Patient John Doe lives at 123 Main St, Springfield.",
    "Contact: jane.doe@example.com, phone 555‑1234."
]

# Extract PII with default batch_size of 8 and confidence threshold of 0.5

result = process_batch(texts, operation="extract_pii")

print(result.summary())
for item in result.get_successful_results():
    print(f"{item.id}: {item.result.entities}")

Customizing Batch Size and Confidence Thresholds

For production workloads requiring strict PII detection or specific throughput targets, instantiate BatchProcessor directly with custom parameters.

from openmed.processing.batch import BatchProcessor

def progress(current: int, total: int, item_result):
    print(f"[{current}/{total}] {item_result.id} – "
          f"{'OK' if item_result.success else 'FAIL'}")

processor = BatchProcessor(
    model_name="pii_en_small",
    operation="extract_pii",
    batch_size=4,                # Process 4 texts per chunk

    confidence_threshold=0.7,    # Stricter confidence cutoff

    continue_on_error=False,     # Stop on first failure

)

texts = ["Patient record 1...", "Patient record 2..."]  # Your clinical notes

result = processor.process_texts(texts, progress_callback=progress)

The progress_callback receives the cumulative count, total items, and the BatchItemResult for real-time monitoring.

Processing Files from Directories

To batch-process entire directories of clinical notes, use process_directory, which globs files, reads UTF-8 content, and streams results through the same PII pipeline (lines 78-86).

from openmed import BatchProcessor

processor = BatchProcessor(operation="extract_pii", batch_size=5)

batch_result = processor.process_directory(
    directory="data/clinical_notes",
    pattern="*.txt",
    recursive=True,
    progress_callback=lambda cur, tot, r: print(f"{cur}/{tot} – {r.id}")
)

print(f"Successfully extracted PII from {batch_result.successful_items} files.")

Streaming Large Datasets with Iterators

For memory-constrained environments processing millions of records, use iter_process (lines 124-131) to yield results one at a time without materializing the entire batch in memory.

from openmed.processing.batch import BatchProcessor

processor = BatchProcessor(operation="extract_pii", batch_size=10)

for item_result in processor.iter_process(very_large_text_sequence):
    if item_result.success:
        # Process item_result.result (PredictionResult)

        entities = item_result.result.entities
        ...
    else:
        # Handle or log the specific failure

        print(f"Failed to process {item_result.id}: {item_result.error}")

Summary

Batch PII extraction in OpenMed leverages a sophisticated pipeline that maximizes throughput while ensuring accuracy:

  • Efficient chunking via _iter_chunks processes data in configurable batches to balance memory usage and speed.
  • Model reuse ensures the privacy-filter pipeline loads only once per batch, eliminating redundant initialization overhead.
  • Smart merging combines overlapping entities automatically using _apply_pii_smart_merging before final validation.
  • Flexible entry points include the high-level process_batch function and the configurable BatchProcessor class for directory and iterator-based workflows.
  • Robust error handling via continue_on_error allows pipelines to complete partially even when individual texts fail processing.

Frequently Asked Questions

What is the default batch size for BatchProcessor?

The default batch_size is 8 texts per chunk. You can override this in the BatchProcessor constructor or via the process_batch helper. The system validates this value using validate_batch_size in openmed/utils/validation.py (lines 158-174) to prevent memory issues.

How does BatchProcessor handle PII model loading?

The processor reuses a single cached privacy-filter pipeline across the entire batch, as implemented in _extract_pii_batch within openmed/core/pii.py. This prevents the expensive operation of reloading transformer models for every text chunk, significantly improving throughput for large datasets.

Can I process files recursively from subdirectories?

Yes. The process_directory method accepts a recursive=True parameter that globs all matching files in subdirectories. It reads each file as UTF-8 text and processes them through the standard batch PII pipeline, returning aggregated results in a BatchResult object.

What happens if one text fails during batch processing?

Behavior depends on the continue_on_error parameter. When set to True (default), the processor logs the failure via _process_single_item and continues with remaining texts, including partial successes in the final BatchResult. When False, the entire batch operation halts immediately upon the first failure.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →