# How to Perform Batch Processing of Clinical Documents with OpenMED's BatchProcessor

> Learn how to perform batch processing of clinical documents with OpenMED's BatchProcessor. Efficiently extract and de-identify PII across thousands of documents using configurable batches and cached ML pipelines.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-12

---

**The OpenMED `BatchProcessor` class orchestrates high-throughput PII extraction and de-identification across thousands of clinical documents by chunking inputs into configurable batches, caching the underlying ML pipeline, and aggregating results into type-safe data transfer objects.**

The `maziyarpanahi/openmed` library provides a production-ready `BatchProcessor` specifically designed to handle sensitive clinical text at scale. This component sits atop the core MLX inference engine and provides fault-tolerant, batched operations for privacy filtering tasks. Understanding how to configure its parameters and leverage its internal caching mechanisms is essential for optimizing throughput in healthcare NLP pipelines.

## Core Architecture and Data Flow

The batch processing system relies on a hierarchy of data classes defined in [[`openmed/processing/batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/batch.py)](https://github.com/maziyarpanahi/openmed/blob/master/openmed/processing/batch.py).

### Primary Data Structures

- **`BatchItem`** (lines ≈ 60‑80): A dataclass that encapsulates a single document, storing `id`, `text`, and optional `metadata` dictionaries. This abstraction allows the processor to handle raw strings or rich objects carrying patient IDs or source system tags.

- **`BatchItemResult`** (lines ≈ 90‑120): Captures the outcome for an individual document, including the processed `result` (e.g., redacted text or entity list), an `error` string if processing failed, and per-item latency.

- **`BatchResult`** (lines ≈ 130‑170): Aggregates a list of `BatchItemResult` objects, tracking `total_processing_time` and exposing a human-readable `summary` property for quick logging.

### Pipeline Caching and Validation

Before processing begins, the constructor invokes `validate_batch_size` from [[`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py)](https://github.com/maziyarpanahi/openmed/blob/master/openmed/utils/validation.py#L158) to ensure the requested batch size is an integer greater than zero and below library-defined limits.

To minimize cold-start latency, the processor implements lazy loading via `_get_shared_loader`. This method caches the underlying MLX pipeline (e.g., the privacy filter from [[`openmed/mlx/models/privacy_filter.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/models/privacy_filter.py)](https://github.com/maziyarpanahi/openmed/blob/master/openmed/mlx/models/privacy_filter.py)) so that model weights are loaded only once per job and reused across all batches, a behavior verified in `test_batch_processor_caches_one_privacy-filter_pipeline_per_job`.

## Configuring Batch Processing Parameters

The `BatchProcessor` constructor accepts several critical parameters that control behavior and output quality.

- **`operation`**: Specifies the transformation to apply. Valid values include `"extract_pii"` (entity recognition), `"redact"` (masking), and `"deidentify"` (placeholder replacement). Each operation type sets an implicit default confidence threshold.

- **`batch_size`**: The number of documents processed per inference call. This value is validated strictly; invalid inputs raise a `ValidationError` before any GPU memory is allocated.

- **`confidence_threshold`**: Filters model predictions. Extraction operations default to `0.5`, while de-identification uses a stricter default of `0.7` to prevent accidental leakage of sensitive tokens.

- **`continue_on_error`**: A boolean flag that determines whether the entire batch run should abort on a single failure or record the error in the `BatchItemResult.error` field and proceed.

## Implementation Examples

### Basic PII Extraction from Raw Text

For straightforward use cases, pass a list of raw strings to the `process_batch` method. This pattern relies on the default `confidence_threshold` of `0.5` for extraction tasks.

```python
from openmed.processing.batch import BatchProcessor

# Prepare clinical documents

documents = [
    "Patient John Doe (MRN: 12345) presented with chest pain.",
    "The MRI showed a 2 cm lesion in the left temporal lobe.",
    "Prescribed amoxicillin 500 mg PO BID for 7 days."
]

# Initialize processor for entity extraction

processor = BatchProcessor(
    operation="extract_pii",
    batch_size=2,
    confidence_threshold=0.5
)

# Execute batch processing

result = processor.process_batch(documents)

# Inspect aggregated results

print(result.summary)
for item in result.items:
    if item.error:
        print(f"[{item.id}] Error: {item.error}")
    else:
        print(f"[{item.id}] Entities: {item.result}")

```

The `process_batch` method (lines ≈ 150‑200 in [`batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/batch.py)) wraps input strings into `BatchItem` instances automatically, forwards them to `_get_analyze_text` for inference, and returns a populated `BatchResult`.

### Advanced De-identification with Metadata and Error Resilience

When processing documents that require traceability, instantiate `BatchItem` objects directly to attach metadata. Enable `continue_on_error` to ensure partial failures do not invalidate the entire workload.

```python
from openmed.processing.batch import BatchProcessor, BatchItem

# Create items with external identifiers

items = [
    BatchItem(id="doc-001", text="Patient A...", metadata={"patient_id": "P001"}),
    BatchItem(id="doc-002", text="Patient B...", metadata={"patient_id": "P002"}),
]

# Configure for de-identification with strict confidence

processor = BatchProcessor(
    operation="deidentify",
    batch_size=10,
    confidence_threshold=0.7,
    continue_on_error=True
)

# Process batch

batch_result = processor.process_batch(items)

# Collect successful de-identified texts

deidentified = {
    item.id: item.result 
    for item in batch_result.items 
    if not item.error
}

```

This pattern leverages the `metadata` field to preserve source system context while the `BatchProcessor` handles the low-level MLX interaction.

### Monitoring Performance with BatchMetrics

To debug throughput bottlenecks, instrument your pipeline with the `BatchMetrics` helper located at line 335 of [[`openmed/utils/profiling.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/profiling.py)](https://github.com/maziyarpanahi/openmed/blob/master/openmed/utils/profiling.py#L335).

```python
from openmed.utils.profiling import BatchMetrics

# Assuming batch_result from previous examples

metrics = BatchMetrics(
    items=batch_result.items, 
    total_time_ms=batch_result.total_processing_time
)

# Render markdown table of latencies

print(metrics.to_markdown())

```

The `BatchMetrics` class calculates per-item and cumulative latency, outputting a formatted table suitable for logging or dashboards.

## Key Source Files Reference

| Component | File | Purpose |
|-----------|------|---------|
| **BatchProcessor** | [`openmed/processing/batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/batch.py) | Core orchestration logic, DTO definitions, and the `process_batch` entry point. |
| **Input Validation** | [`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py) | Contains `validate_batch_size` (line 158) for parameter sanitization. |
| **Profiling** | [`openmed/utils/profiling.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/profiling.py) | Implements `BatchMetrics` (line 335) for latency analysis. |
| **MLX Pipeline** | [`openmed/mlx/models/privacy_filter.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/mlx/models/privacy_filter.py) | Privacy filter model used for PII extraction and redaction. |
| **Usage Examples** | [`examples/pii_batch_processing.py`](https://github.com/maziyarpanahi/openmed/blob/main/examples/pii_batch_processing.py) | End-to-end demonstration script. |
| **Unit Tests** | [`tests/unit/test_batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_batch.py) | Verifies caching behavior, error handling, and default thresholds. |

## Summary

- The **`BatchProcessor`** in [`openmed/processing/batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/batch.py) is the primary entry point for bulk clinical document processing, handling batching, caching, and error aggregation automatically.
- **Data classes** (`BatchItem`, `BatchItemResult`, `BatchResult`) provide type-safe containers for documents and results, supporting optional metadata and detailed error tracking.
- **Configuration** is controlled via `operation`, `batch_size`, `confidence_threshold`, and `continue_on_error`, with strict validation enforced by [`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py).
- **Pipeline caching** via `_get_shared_loader` ensures that model weights are loaded only once per job, significantly reducing latency for large corpora.
- **Performance monitoring** is available through the `BatchMetrics` class in [`openmed/utils/profiling.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/profiling.py), which renders detailed latency tables for optimization.

## Frequently Asked Questions

### What is the maximum batch size supported by OpenMED's BatchProcessor?

The library enforces an upper bound on `batch_size` to prevent GPU memory exhaustion. While the exact limit depends on the model configuration and hardware, the `validate_batch_size` function in [`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py) (line 158) ensures the value is a positive integer below the library-wide maximum. Attempting to initialize `BatchProcessor` with an excessive batch size raises a `ValidationError` before processing begins.

### How does the BatchProcessor handle documents that fail during processing?

When the `continue_on_error` parameter is set to `True`, the processor catches exceptions during the `_get_analyze_text` phase and stores the error message in the `BatchItemResult.error` field. The remaining documents in the batch continue processing normally. If `continue_on_error` is `False`, the first exception propagates and halts the entire batch operation. This behavior is tested in [`tests/unit/test_batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/unit/test_batch.py) under the `test_batch_processor_continue_on_error` scenario.

### Can I use custom metadata to track documents through the pipeline?

Yes. By constructing `BatchItem` objects manually instead of passing raw strings, you can populate the `metadata` dictionary with custom key-value pairs such as patient IDs or source filenames. This metadata is preserved in the resulting `BatchItemResult` and can be used to correlate outputs with upstream database records without modifying the clinical text content.

### What confidence threshold should I set for de-identification versus extraction?

The `BatchProcessor` applies operation-specific defaults: **0.5** for `extract_pii` to ensure high recall of potential entities, and **0.7** for `deidentify` to maximize precision and prevent accidental retention of sensitive tokens. You can override these defaults via the `confidence_threshold` constructor argument. For compliance-critical deployments, use the stricter 0.7 threshold or higher to minimize false negatives in redaction tasks.