# How to Implement Batch PII Extraction with BatchProcessor in OpenMed

> Learn how to implement batch PII extraction efficiently with OpenMed's BatchProcessor. Discover automatic model caching, entity merging, and error aggregation for large text collections.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**OpenMed's `BatchProcessor` class enables efficient, chunk-wise PII extraction from large text collections by automatically handling model caching, smart entity merging, and error aggregation when configured with the `"extract_pii"` operation.**

Batch PII extraction is essential for processing clinical notes and patient records at scale without exhausting memory or reloading models unnecessarily. The **OpenMed** library accelerates this workflow through a dedicated `BatchProcessor` that orchestrates the entire pipeline from input validation to result aggregation. This guide demonstrates how to implement batch PII extraction with BatchProcessor using the actual source implementation from the `maziyarpanahi/openmed` repository.

## Understanding the Batch PII Architecture

### The Core Extraction Pipeline

When the `BatchProcessor` operation is set to `"extract_pii"`, it delegates heavy lifting to the private helper `_extract_pii_batch` defined in [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py) (lines 30-34). According to the OpenMed source code, this helper executes five critical steps:

1. **Input normalization** – Optionally strips accents and resolves the language-specific model via `_resolve_effective_pii_model` (lines 51-64).
2. **Pipeline reuse** – Maintains a single cached privacy-filter pipeline instance so models load only once per batch rather than per item.
3. **Backend execution** – Calls `create_privacy_filter_pipeline` (for privacy-filter models) or the standard `analyze_text` function on the entire text list.
4. **Smart merging** – Applies `_apply_pii_smart_merging` (lines 84-90) to combine overlapping entity spans into coherent PII entries.
5. **Validation** – Runs `validate_entity_spans` (lines 126-130) to ensure extracted boundaries align with actual text content before returning results.

### Chunking and Orchestration

The `BatchProcessor` class in [`openmed/processing/batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/batch.py) manages the high-level workflow:

- **Initialization** – Validates `batch_size` using `validate_batch_size` from [`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py) (lines 158-174).
- **Chunking** – Splits input lists into slices via `_iter_chunks` (lines 99-103).
- **Per-chunk processing** – Routes PII-specific logic through `_process_pii_chunk` (lines 98-102).
- **Result aggregation** – Constructs a `BatchResult` object (lines 70-106) containing per-item results, execution timings, and human-readable summaries.
- **Error handling** – Implements `continue_on_error` logic via `_process_single_item` (lines 68-72) to either halt or log failures.

## Implementation Examples

### Basic Batch Processing with process_batch

For rapid implementation, use the `process_batch` convenience function, which instantiates a `BatchProcessor` behind the scenes and invokes `processor.process_texts()` (lines 146-155).

```python
from openmed import process_batch

texts = [
    "Patient John Doe lives at 123 Main St, Springfield.",
    "Contact: jane.doe@example.com, phone 555‑1234."
]

# Extract PII with default batch_size of 8 and confidence threshold of 0.5

result = process_batch(texts, operation="extract_pii")

print(result.summary())
for item in result.get_successful_results():
    print(f"{item.id}: {item.result.entities}")

```

### Customizing Batch Size and Confidence Thresholds

For production workloads requiring strict PII detection or specific throughput targets, instantiate `BatchProcessor` directly with custom parameters.

```python
from openmed.processing.batch import BatchProcessor

def progress(current: int, total: int, item_result):
    print(f"[{current}/{total}] {item_result.id} – "
          f"{'OK' if item_result.success else 'FAIL'}")

processor = BatchProcessor(
    model_name="pii_en_small",
    operation="extract_pii",
    batch_size=4,                # Process 4 texts per chunk

    confidence_threshold=0.7,    # Stricter confidence cutoff

    continue_on_error=False,     # Stop on first failure

)

texts = ["Patient record 1...", "Patient record 2..."]  # Your clinical notes

result = processor.process_texts(texts, progress_callback=progress)

```

The `progress_callback` receives the cumulative count, total items, and the `BatchItemResult` for real-time monitoring.

### Processing Files from Directories

To batch-process entire directories of clinical notes, use `process_directory`, which globs files, reads UTF-8 content, and streams results through the same PII pipeline (lines 78-86).

```python
from openmed import BatchProcessor

processor = BatchProcessor(operation="extract_pii", batch_size=5)

batch_result = processor.process_directory(
    directory="data/clinical_notes",
    pattern="*.txt",
    recursive=True,
    progress_callback=lambda cur, tot, r: print(f"{cur}/{tot} – {r.id}")
)

print(f"Successfully extracted PII from {batch_result.successful_items} files.")

```

### Streaming Large Datasets with Iterators

For memory-constrained environments processing millions of records, use `iter_process` (lines 124-131) to yield results one at a time without materializing the entire batch in memory.

```python
from openmed.processing.batch import BatchProcessor

processor = BatchProcessor(operation="extract_pii", batch_size=10)

for item_result in processor.iter_process(very_large_text_sequence):
    if item_result.success:
        # Process item_result.result (PredictionResult)

        entities = item_result.result.entities
        ...
    else:
        # Handle or log the specific failure

        print(f"Failed to process {item_result.id}: {item_result.error}")

```

## Summary

Batch PII extraction in OpenMed leverages a sophisticated pipeline that maximizes throughput while ensuring accuracy:

- **Efficient chunking** via `_iter_chunks` processes data in configurable batches to balance memory usage and speed.
- **Model reuse** ensures the privacy-filter pipeline loads only once per batch, eliminating redundant initialization overhead.
- **Smart merging** combines overlapping entities automatically using `_apply_pii_smart_merging` before final validation.
- **Flexible entry points** include the high-level `process_batch` function and the configurable `BatchProcessor` class for directory and iterator-based workflows.
- **Robust error handling** via `continue_on_error` allows pipelines to complete partially even when individual texts fail processing.

## Frequently Asked Questions

### What is the default batch size for BatchProcessor?

The default `batch_size` is **8** texts per chunk. You can override this in the `BatchProcessor` constructor or via the `process_batch` helper. The system validates this value using `validate_batch_size` in [`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py) (lines 158-174) to prevent memory issues.

### How does BatchProcessor handle PII model loading?

The processor reuses a single cached privacy-filter pipeline across the entire batch, as implemented in `_extract_pii_batch` within [`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py). This prevents the expensive operation of reloading transformer models for every text chunk, significantly improving throughput for large datasets.

### Can I process files recursively from subdirectories?

Yes. The `process_directory` method accepts a `recursive=True` parameter that globs all matching files in subdirectories. It reads each file as UTF-8 text and processes them through the standard batch PII pipeline, returning aggregated results in a `BatchResult` object.

### What happens if one text fails during batch processing?

Behavior depends on the `continue_on_error` parameter. When set to `True` (default), the processor logs the failure via `_process_single_item` and continues with remaining texts, including partial successes in the final `BatchResult`. When `False`, the entire batch operation halts immediately upon the first failure.