How to Perform Batch Processing of Clinical Documents with OpenMED's BatchProcessor
The OpenMED BatchProcessor class orchestrates high-throughput PII extraction and de-identification across thousands of clinical documents by chunking inputs into configurable batches, caching the underlying ML pipeline, and aggregating results into type-safe data transfer objects.
The maziyarpanahi/openmed library provides a production-ready BatchProcessor specifically designed to handle sensitive clinical text at scale. This component sits atop the core MLX inference engine and provides fault-tolerant, batched operations for privacy filtering tasks. Understanding how to configure its parameters and leverage its internal caching mechanisms is essential for optimizing throughput in healthcare NLP pipelines.
Core Architecture and Data Flow
The batch processing system relies on a hierarchy of data classes defined in [openmed/processing/batch.py](https://github.com/maziyarpanahi/openmed/blob/master/openmed/processing/batch.py).
Primary Data Structures
-
BatchItem(lines ≈ 60‑80): A dataclass that encapsulates a single document, storingid,text, and optionalmetadatadictionaries. This abstraction allows the processor to handle raw strings or rich objects carrying patient IDs or source system tags. -
BatchItemResult(lines ≈ 90‑120): Captures the outcome for an individual document, including the processedresult(e.g., redacted text or entity list), anerrorstring if processing failed, and per-item latency. -
BatchResult(lines ≈ 130‑170): Aggregates a list ofBatchItemResultobjects, trackingtotal_processing_timeand exposing a human-readablesummaryproperty for quick logging.
Pipeline Caching and Validation
Before processing begins, the constructor invokes validate_batch_size from [openmed/utils/validation.py](https://github.com/maziyarpanahi/openmed/blob/master/openmed/utils/validation.py#L158) to ensure the requested batch size is an integer greater than zero and below library-defined limits.
To minimize cold-start latency, the processor implements lazy loading via _get_shared_loader. This method caches the underlying MLX pipeline (e.g., the privacy filter from [openmed/mlx/models/privacy_filter.py](https://github.com/maziyarpanahi/openmed/blob/master/openmed/mlx/models/privacy_filter.py)) so that model weights are loaded only once per job and reused across all batches, a behavior verified in test_batch_processor_caches_one_privacy-filter_pipeline_per_job.
Configuring Batch Processing Parameters
The BatchProcessor constructor accepts several critical parameters that control behavior and output quality.
-
operation: Specifies the transformation to apply. Valid values include"extract_pii"(entity recognition),"redact"(masking), and"deidentify"(placeholder replacement). Each operation type sets an implicit default confidence threshold. -
batch_size: The number of documents processed per inference call. This value is validated strictly; invalid inputs raise aValidationErrorbefore any GPU memory is allocated. -
confidence_threshold: Filters model predictions. Extraction operations default to0.5, while de-identification uses a stricter default of0.7to prevent accidental leakage of sensitive tokens. -
continue_on_error: A boolean flag that determines whether the entire batch run should abort on a single failure or record the error in theBatchItemResult.errorfield and proceed.
Implementation Examples
Basic PII Extraction from Raw Text
For straightforward use cases, pass a list of raw strings to the process_batch method. This pattern relies on the default confidence_threshold of 0.5 for extraction tasks.
from openmed.processing.batch import BatchProcessor
# Prepare clinical documents
documents = [
"Patient John Doe (MRN: 12345) presented with chest pain.",
"The MRI showed a 2 cm lesion in the left temporal lobe.",
"Prescribed amoxicillin 500 mg PO BID for 7 days."
]
# Initialize processor for entity extraction
processor = BatchProcessor(
operation="extract_pii",
batch_size=2,
confidence_threshold=0.5
)
# Execute batch processing
result = processor.process_batch(documents)
# Inspect aggregated results
print(result.summary)
for item in result.items:
if item.error:
print(f"[{item.id}] Error: {item.error}")
else:
print(f"[{item.id}] Entities: {item.result}")
The process_batch method (lines ≈ 150‑200 in batch.py) wraps input strings into BatchItem instances automatically, forwards them to _get_analyze_text for inference, and returns a populated BatchResult.
Advanced De-identification with Metadata and Error Resilience
When processing documents that require traceability, instantiate BatchItem objects directly to attach metadata. Enable continue_on_error to ensure partial failures do not invalidate the entire workload.
from openmed.processing.batch import BatchProcessor, BatchItem
# Create items with external identifiers
items = [
BatchItem(id="doc-001", text="Patient A...", metadata={"patient_id": "P001"}),
BatchItem(id="doc-002", text="Patient B...", metadata={"patient_id": "P002"}),
]
# Configure for de-identification with strict confidence
processor = BatchProcessor(
operation="deidentify",
batch_size=10,
confidence_threshold=0.7,
continue_on_error=True
)
# Process batch
batch_result = processor.process_batch(items)
# Collect successful de-identified texts
deidentified = {
item.id: item.result
for item in batch_result.items
if not item.error
}
This pattern leverages the metadata field to preserve source system context while the BatchProcessor handles the low-level MLX interaction.
Monitoring Performance with BatchMetrics
To debug throughput bottlenecks, instrument your pipeline with the BatchMetrics helper located at line 335 of [openmed/utils/profiling.py](https://github.com/maziyarpanahi/openmed/blob/master/openmed/utils/profiling.py#L335).
from openmed.utils.profiling import BatchMetrics
# Assuming batch_result from previous examples
metrics = BatchMetrics(
items=batch_result.items,
total_time_ms=batch_result.total_processing_time
)
# Render markdown table of latencies
print(metrics.to_markdown())
The BatchMetrics class calculates per-item and cumulative latency, outputting a formatted table suitable for logging or dashboards.
Key Source Files Reference
| Component | File | Purpose |
|---|---|---|
| BatchProcessor | openmed/processing/batch.py |
Core orchestration logic, DTO definitions, and the process_batch entry point. |
| Input Validation | openmed/utils/validation.py |
Contains validate_batch_size (line 158) for parameter sanitization. |
| Profiling | openmed/utils/profiling.py |
Implements BatchMetrics (line 335) for latency analysis. |
| MLX Pipeline | openmed/mlx/models/privacy_filter.py |
Privacy filter model used for PII extraction and redaction. |
| Usage Examples | examples/pii_batch_processing.py |
End-to-end demonstration script. |
| Unit Tests | tests/unit/test_batch.py |
Verifies caching behavior, error handling, and default thresholds. |
Summary
- The
BatchProcessorinopenmed/processing/batch.pyis the primary entry point for bulk clinical document processing, handling batching, caching, and error aggregation automatically. - Data classes (
BatchItem,BatchItemResult,BatchResult) provide type-safe containers for documents and results, supporting optional metadata and detailed error tracking. - Configuration is controlled via
operation,batch_size,confidence_threshold, andcontinue_on_error, with strict validation enforced byopenmed/utils/validation.py. - Pipeline caching via
_get_shared_loaderensures that model weights are loaded only once per job, significantly reducing latency for large corpora. - Performance monitoring is available through the
BatchMetricsclass inopenmed/utils/profiling.py, which renders detailed latency tables for optimization.
Frequently Asked Questions
What is the maximum batch size supported by OpenMED's BatchProcessor?
The library enforces an upper bound on batch_size to prevent GPU memory exhaustion. While the exact limit depends on the model configuration and hardware, the validate_batch_size function in openmed/utils/validation.py (line 158) ensures the value is a positive integer below the library-wide maximum. Attempting to initialize BatchProcessor with an excessive batch size raises a ValidationError before processing begins.
How does the BatchProcessor handle documents that fail during processing?
When the continue_on_error parameter is set to True, the processor catches exceptions during the _get_analyze_text phase and stores the error message in the BatchItemResult.error field. The remaining documents in the batch continue processing normally. If continue_on_error is False, the first exception propagates and halts the entire batch operation. This behavior is tested in tests/unit/test_batch.py under the test_batch_processor_continue_on_error scenario.
Can I use custom metadata to track documents through the pipeline?
Yes. By constructing BatchItem objects manually instead of passing raw strings, you can populate the metadata dictionary with custom key-value pairs such as patient IDs or source filenames. This metadata is preserved in the resulting BatchItemResult and can be used to correlate outputs with upstream database records without modifying the clinical text content.
What confidence threshold should I set for de-identification versus extraction?
The BatchProcessor applies operation-specific defaults: 0.5 for extract_pii to ensure high recall of potential entities, and 0.7 for deidentify to maximize precision and prevent accidental retention of sensitive tokens. You can override these defaults via the confidence_threshold constructor argument. For compliance-critical deployments, use the stricter 0.7 threshold or higher to minimize false negatives in redaction tasks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →