How to Customize OpenMed Output to CSV Format: A Complete Guide
OpenMed provides built-in CSV export capabilities through the OutputFormatter class in openmed/processing/outputs.py, which converts structured prediction results into flat dictionary rows via the to_csv_rows method.
Customizing OpenMed output to CSV format enables seamless integration of medical entity recognition results into data analysis pipelines and reporting workflows. The maziyarpanahi/openmed repository includes a dedicated formatting module that handles the conversion from raw model predictions to standardized CSV rows. By leveraging the format_predictions helper and OutputFormatter class, you can extract clinical entities with their metadata, confidence scores, and positional information in a tabular format compatible with pandas, Excel, or SQL databases.
Understanding the OpenMed Output Pipeline
Before customizing CSV output, you need to understand how OpenMed structures prediction results. The pipeline transforms raw model outputs into structured objects before flattening them to CSV rows.
The PredictionResult Structure
According to the source code in openmed/processing/outputs.py, every prediction generates a PredictionResult object (lines 77+) that encapsulates:
- The original input text
- Model name and processing timestamp
- A list of
EntityPredictionobjects containing entity text, label, confidence scores, and character spans
This structure serves as the intermediate representation between raw transformer outputs and final CSV rows.
The OutputFormatter Class
The OutputFormatter class (found in openmed/processing/outputs.py) normalizes raw predictions and handles format-specific conversions. Its core responsibilities include:
- Fixing token spans to align with the original text
- Filtering entities by confidence threshold
- Grouping adjacent entities when configured
- Converting structured data to CSV-compatible dictionaries via
to_csv_rows
Converting OpenMed Predictions to CSV Format
OpenMed offers two primary approaches for CSV generation: using the high-level format_predictions wrapper or invoking the OutputFormatter directly for fine-grained control.
Using the format_predictions Helper
The format_predictions function (lines 551+ in openmed/processing/outputs.py) provides a convenient one-line interface. When you specify output_format="csv", the function automatically instantiates an OutputFormatter and routes results through the to_csv_rows method.
Direct CSV Row Generation with OutputFormatter
For production pipelines requiring custom processing, instantiate OutputFormatter directly and call to_csv_rows on a PredictionResult object. This approach allows you to intercept results from the OpenMedService class before serialization.
Practical Implementation Examples
The following code examples demonstrate complete workflows for customizing OpenMed CSV output in different scenarios.
Basic CSV Export from Raw Predictions
Convert raw NER model outputs to CSV-ready dictionaries using the format_predictions helper:
from openmed.processing.outputs import format_predictions
# Simulated raw predictions from a NER model
raw_predictions = [
{"word": "Aspirin", "score": 0.96, "start": 10, "end": 17, "entity": "B-MEDICATION"},
{"word": "headache", "score": 0.88, "start": 22, "end": 30, "entity": "B-CONDITION"},
]
original_text = "Patient was prescribed Aspirin for headache."
model_name = "my-ner-model"
# Request CSV-compatible rows directly
csv_rows = format_predictions(
predictions=raw_predictions,
original_text=original_text,
model_name=model_name,
output_format="csv", # <-- triggers to_csv_rows
include_confidence=True, # optional, passed to OutputFormatter
confidence_threshold=0.5, # optional filter
)
print(csv_rows)
The output contains a list of dictionaries with standardized keys: text, label, confidence, start, end, model_name, timestamp, processing_time, and original_text.
Writing to a CSV File
Persist the formatted rows to disk using Python's built-in csv module:
import csv
from openmed.processing.outputs import format_predictions
# (same raw_predictions / original_text as above)
csv_rows = format_predictions(
predictions=raw_predictions,
original_text=original_text,
model_name=model_name,
output_format="csv",
)
fieldnames = [
"text", "label", "confidence", "start", "end",
"model_name", "timestamp", "processing_time", "original_text"
]
with open("ner_output.csv", "w", newline="", encoding="utf-8") as fp:
writer = csv.DictWriter(fp, fieldnames=fieldnames, extrasaction="ignore")
writer.writeheader()
writer.writerows(csv_rows)
The resulting ner_output.csv contains one line per detected entity with complete metadata for downstream analytics.
Adding Custom Metadata to CSV Rows
OpenMed automatically flattens metadata dictionaries using the metadata_<key> naming convention. Attach custom fields to raw predictions for extended CSV output:
raw_predictions = [
{
"word": "Aspirin",
"score": 0.96,
"start": 10,
"end": 17,
"entity": "B-MEDICATION",
"metadata": {"sentence_index": 0, "source": "clinical_note"}
},
]
csv_rows = format_predictions(
predictions=raw_predictions,
original_text=original_text,
model_name=model_name,
output_format="csv",
)
# csv_rows now contain "metadata_sentence_index" and "metadata_source" fields
Integrating CSV Export in Production Pipelines
When processing text through OpenMed's service layer, extract CSV rows directly from the PredictionResult object:
from openmed.service.app import OpenMedService
from openmed.processing.outputs import OutputFormatter
service = OpenMedService()
result = service.process_text("Patient was prescribed Aspirin for headache.")
formatter = OutputFormatter()
csv_rows = formatter.to_csv_rows(result) # direct call on the result object
# Save or stream csv_rows as needed
Summary
- OpenMed CSV customization centers on the
OutputFormatterclass inopenmed/processing/outputs.py. - The
to_csv_rowsmethod convertsPredictionResultobjects into flat dictionaries suitable for CSV serialization. - Use
format_predictionswithoutput_format="csv"for quick one-off conversions, or instantiateOutputFormatterdirectly for pipeline integration. - Custom metadata automatically appears in CSV output with the
metadata_prefix. - The CSV format includes entity text, labels, confidence scores, character spans, model metadata, and timestamps.
Frequently Asked Questions
What fields are included in OpenMed CSV output?
The to_csv_rows method in openmed/processing/outputs.py generates dictionaries containing: text (the entity text), label (the entity type), confidence (the prediction score), start and end (character positions), model_name (the source model), timestamp (processing time), processing_time (duration), and original_text (the full input). Custom metadata fields appear with metadata_ prefixes.
How do I filter low-confidence predictions before CSV export?
Pass the confidence_threshold parameter to format_predictions or OutputFormatter.format_predictions. Values below this threshold (default 0.5) are excluded from the PredictionResult.entities list before CSV conversion occurs, ensuring only high-quality predictions appear in your output file.
Can I customize the column names in the OpenMed CSV output?
OpenMed generates standardized field names in to_csv_rows. To use custom column names, transform the returned dictionaries before writing to CSV. Iterate through csv_rows and remap keys using a dictionary comprehension, or use csv.DictWriter with the extrasaction="ignore" parameter and specify your preferred fieldnames order.
Where is the CSV formatting logic implemented in OpenMed?
The core logic resides in openmed/processing/outputs.py at lines 515+ (the to_csv_rows method) and lines 551+ (the format_predictions wrapper). The OutputFormatter class handles the conversion pipeline, while EntityPrediction (lines 55+) and PredictionResult (lines 77+) define the data structures being converted.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →