# Optimizing OpenMed Performance with Batching: A Complete Guide to High-Throughput Clinical NLP

> Boost OpenMed clinical NLP performance with batching. This guide explains how BatchProcessor reuses model loaders and creates GPU-efficient tensors for high throughput.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: performance
- Published: 2026-06-12

---

**OpenMed maximizes clinical NLP throughput by processing documents in batches through the `BatchProcessor` class, which reuses model loaders and aggregates texts into GPU-efficient tensors to eliminate per-document initialization overhead.**

OpenMed, maintained in the `maziyarpanahi/openmed` repository, provides enterprise-grade batching infrastructure for processing large-scale clinical text corpora. The batching system implemented in [`openmed/processing/batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/batch.py) transforms sequential processing into high-throughput pipelines through intelligent chunking, model loader reuse, and memory-bound execution. By leveraging these batching primitives, you can achieve significant performance gains when running operations like PII extraction, de-identification, or clinical entity recognition across thousands of documents.

## Core Batching Architecture

The batching subsystem centers on the **`BatchProcessor`** class defined at line 52 of [`openmed/processing/batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/batch.py). This orchestrator manages the entire batch lifecycle, from creating **`BatchItem`** objects through result aggregation, while maintaining strict memory boundaries and resource efficiency.

### Chunking Strategy and Memory Control

The **`_iter_chunks`** method (lines 99-103) divides input lists into contiguous sub-lists of size `batch_size`. This chunking prevents memory exhaustion by ensuring only a fixed number of texts reside in GPU memory at any moment, making it safe to process corpora larger than available RAM.

### Model Loader Reuse

Instead of initializing models for every document, the **`_get_shared_loader`** method (lines 53-67) creates a single **`ModelLoader`** instance that persists across all chunks. This reuse eliminates the costly overhead of repeatedly loading transformer weights into GPU memory, which typically dominates processing time in naive sequential implementations.

### Privacy Pipeline Caching

For PII-related operations (`extract_pii` and `deidentify`), the **`_get_privacy_filter_pipeline`** method (lines 69-97) instantiates one privacy-filter pipeline per effective model configuration. This pipeline caches tokenizer and model instances, ensuring that sensitive data processing benefits from the same reuse patterns as standard analysis tasks.

## Performance Optimization Mechanisms

Batching improves OpenMed performance through four critical optimization paths:

- **Reduced initialization overhead** – Model loading occurs once via `_get_shared_loader` rather than once per document, eliminating redundant disk I/O and GPU memory allocation.
- **GPU parallelism** – By collating texts into fixed-size tensors, the underlying HuggingFace transformers pipeline executes a single forward pass for the entire batch rather than sequential passes, maximizing GPU utilization.
- **Cache-friendly I/O** – When processing files through `process_directory`, the system reads each file exactly once and immediately queues its content for chunking, minimizing disk seek operations.
- **Bounded memory usage** – The `batch_size` parameter acts as a hard memory cap, preventing out-of-memory errors when processing variable-length clinical narratives on resource-constrained hardware.

## Implementing Batch Processing

### Basic Text Batching

The **`process_batch`** function (lines 46-48) provides the simplest entry point for text-only workloads. It wraps `BatchProcessor` configuration and returns a **`BatchResult`** with built-in summary statistics.

```python
from openmed import process_batch

texts = [
    "Patient reports fever and cough for three days.",
    "No significant abnormalities on chest X-ray.",
    "History of hypertension; current meds include Lisinopril."
]

result = process_batch(
    texts,
    model_name="disease_detection_superclinical",
    batch_size=4,  # Must not exceed max_batch_size (100 default)

    operation="extract_pii",
    confidence_threshold=0.6,
    continue_on_error=True,
)

print(result.summary())

```

### Advanced BatchProcessor Configuration

For production workflows requiring progress monitoring or custom aggregation strategies, instantiate **`BatchProcessor`** directly and provide a callback function.

```python
from openmed.processing.batch import BatchProcessor

def progress(current, total, item_result):
    print(f"[{current}/{total}] {item_result.id}: "
          f"{'OK' if item_result.success else 'FAIL'}")

processor = BatchProcessor(
    model_name="disease_detection_superclinical",
    operation="analyze_text",
    batch_size=8,
    aggregation_strategy="simple",
    confidence_threshold=0.0,
)

texts = ["..."] * 20  # Large clinical note collection

result = processor.process_texts(texts, progress_callback=progress)

print("Success rate:", result.success_rate)

```

The underlying processing loop in `_process_items` (lines 28-52) invokes this callback after each item completion, enabling real-time monitoring without blocking the batch pipeline.

### Processing Files and Directories

Use **`process_directory`** to recursively scan paths and batch-process matching files without manual text extraction.

```python
from openmed.processing.batch import BatchProcessor

processor = BatchProcessor(
    model_name="disease_detection_superclinical",
    operation="deidentify",
    batch_size=5,
    continue_on_error=False,  # Abort on any file read error

)

result = processor.process_directory(
    directory="clinical_notes/",
    pattern="*.txt",
    recursive=True,
    encoding="utf-8"
)

print(result.summary())

```

This method (lines 78-85) leverages `Path.rglob` for pattern matching and streams file contents directly into the batching pipeline.

### Streaming Large Datasets

For memory-constrained environments processing millions of documents, **`iter_process`** (lines 124-146) yields **`BatchItemResult`** objects one-by-one rather than accumulating all results in memory.

```python
from openmed.processing.batch import BatchProcessor

processor = BatchProcessor(
    model_name="disease_detection_superclinical",
    operation="analyze_text",
    batch_size=10,
)

for item_result in processor.iter_process(very_long_text_list):
    if item_result.success:
        print(item_result.result.to_dict())
    else:
        print("Error:", item_result.error)

```

This iterator pattern maintains constant memory regardless of input corpus size.

## Validation and Error Handling

The **`validate_batch_size`** function in [`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py) (line 54) enforces that batch sizes are positive integers not exceeding the configurable hard limit of 100. This validation prevents resource exhaustion from accidentally oversized batches.

When `continue_on_error=True`, the **`_process_pii_chunk`** method (lines 57-66) implements intelligent fallback: if a chunk fails, the processor retries each item individually using the same operation function, isolating failures without terminating the entire batch. Failed items populate the `BatchResult` with error details while successful items return normally.

## Summary

- **BatchProcessor** in [`openmed/processing/batch.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/batch.py) orchestrates high-throughput clinical NLP by managing chunking, loader reuse, and result aggregation.
- **Chunking** via `_iter_chunks` bounds memory usage by processing fixed-size sub-lists rather than loading entire corpora into GPU memory.
- **Model reuse** through `_get_shared_loader` eliminates per-document initialization overhead, providing the primary performance gain for transformer-based models.
- **Error resilience** with `continue_on_error` enables graceful degradation by retrying failed chunks at the item level without batch termination.
- **Validation** in [`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py) ensures batch parameters remain within safe operational bounds.

## Frequently Asked Questions

### What is the optimal batch size for OpenMed processing?

The optimal batch size depends on your GPU memory and document length, but should remain at or below the validated limit of 100. Start with 8-16 for clinical notes averaging 500 tokens, then increase gradually while monitoring GPU memory utilization via `nvidia-smi` until you approach the hardware ceiling without exceeding it.

### How does OpenMed handle errors during batch processing?

When `continue_on_error=True`, the batch processor catches exceptions at the chunk level and automatically falls back to item-by-item processing for that specific chunk, as implemented in `_process_pii_chunk`. This isolates failures to individual documents while allowing the remainder of the batch to complete, with errors recorded in the `BatchResult` statistics.

### Can I process files directly without loading text into memory first?

Yes, the `process_directory` method handles file I/O internally, reading each file once and streaming its content into the batching pipeline. This avoids loading entire directories into Python memory, making it suitable for processing terabyte-scale corpora from disk.

### Does batching reduce memory usage or only improve speed?

Batching primarily improves speed through model loader reuse and parallel GPU execution, but it also provides **memory control** through explicit `batch_size` limits. Without batching, frameworks often load entire datasets; OpenMed's `_iter_chunks` ensures only the active batch resides in GPU memory, preventing out-of-memory crashes on large datasets.