How to Customize OpenMed Output to CSV Format: A Complete Guide

OpenMed provides built-in CSV export capabilities through the OutputFormatter class in openmed/processing/outputs.py, which converts structured prediction results into flat dictionary rows via the to_csv_rows method.

Customizing OpenMed output to CSV format enables seamless integration of medical entity recognition results into data analysis pipelines and reporting workflows. The maziyarpanahi/openmed repository includes a dedicated formatting module that handles the conversion from raw model predictions to standardized CSV rows. By leveraging the format_predictions helper and OutputFormatter class, you can extract clinical entities with their metadata, confidence scores, and positional information in a tabular format compatible with pandas, Excel, or SQL databases.

Understanding the OpenMed Output Pipeline

Before customizing CSV output, you need to understand how OpenMed structures prediction results. The pipeline transforms raw model outputs into structured objects before flattening them to CSV rows.

The PredictionResult Structure

According to the source code in openmed/processing/outputs.py, every prediction generates a PredictionResult object (lines 77+) that encapsulates:

  • The original input text
  • Model name and processing timestamp
  • A list of EntityPrediction objects containing entity text, label, confidence scores, and character spans

This structure serves as the intermediate representation between raw transformer outputs and final CSV rows.

The OutputFormatter Class

The OutputFormatter class (found in openmed/processing/outputs.py) normalizes raw predictions and handles format-specific conversions. Its core responsibilities include:

  • Fixing token spans to align with the original text
  • Filtering entities by confidence threshold
  • Grouping adjacent entities when configured
  • Converting structured data to CSV-compatible dictionaries via to_csv_rows

Converting OpenMed Predictions to CSV Format

OpenMed offers two primary approaches for CSV generation: using the high-level format_predictions wrapper or invoking the OutputFormatter directly for fine-grained control.

Using the format_predictions Helper

The format_predictions function (lines 551+ in openmed/processing/outputs.py) provides a convenient one-line interface. When you specify output_format="csv", the function automatically instantiates an OutputFormatter and routes results through the to_csv_rows method.

Direct CSV Row Generation with OutputFormatter

For production pipelines requiring custom processing, instantiate OutputFormatter directly and call to_csv_rows on a PredictionResult object. This approach allows you to intercept results from the OpenMedService class before serialization.

Practical Implementation Examples

The following code examples demonstrate complete workflows for customizing OpenMed CSV output in different scenarios.

Basic CSV Export from Raw Predictions

Convert raw NER model outputs to CSV-ready dictionaries using the format_predictions helper:

from openmed.processing.outputs import format_predictions

# Simulated raw predictions from a NER model

raw_predictions = [
    {"word": "Aspirin", "score": 0.96, "start": 10, "end": 17, "entity": "B-MEDICATION"},
    {"word": "headache", "score": 0.88, "start": 22, "end": 30, "entity": "B-CONDITION"},
]

original_text = "Patient was prescribed Aspirin for headache."
model_name = "my-ner-model"

# Request CSV-compatible rows directly

csv_rows = format_predictions(
    predictions=raw_predictions,
    original_text=original_text,
    model_name=model_name,
    output_format="csv",          # <-- triggers to_csv_rows

    include_confidence=True,     # optional, passed to OutputFormatter

    confidence_threshold=0.5,    # optional filter

)

print(csv_rows)

The output contains a list of dictionaries with standardized keys: text, label, confidence, start, end, model_name, timestamp, processing_time, and original_text.

Writing to a CSV File

Persist the formatted rows to disk using Python's built-in csv module:

import csv
from openmed.processing.outputs import format_predictions

# (same raw_predictions / original_text as above)

csv_rows = format_predictions(
    predictions=raw_predictions,
    original_text=original_text,
    model_name=model_name,
    output_format="csv",
)

fieldnames = [
    "text", "label", "confidence", "start", "end",
    "model_name", "timestamp", "processing_time", "original_text"
]

with open("ner_output.csv", "w", newline="", encoding="utf-8") as fp:
    writer = csv.DictWriter(fp, fieldnames=fieldnames, extrasaction="ignore")
    writer.writeheader()
    writer.writerows(csv_rows)

The resulting ner_output.csv contains one line per detected entity with complete metadata for downstream analytics.

Adding Custom Metadata to CSV Rows

OpenMed automatically flattens metadata dictionaries using the metadata_<key> naming convention. Attach custom fields to raw predictions for extended CSV output:

raw_predictions = [
    {
        "word": "Aspirin",
        "score": 0.96,
        "start": 10,
        "end": 17,
        "entity": "B-MEDICATION",
        "metadata": {"sentence_index": 0, "source": "clinical_note"}
    },
]

csv_rows = format_predictions(
    predictions=raw_predictions,
    original_text=original_text,
    model_name=model_name,
    output_format="csv",
)

# csv_rows now contain "metadata_sentence_index" and "metadata_source" fields

Integrating CSV Export in Production Pipelines

When processing text through OpenMed's service layer, extract CSV rows directly from the PredictionResult object:

from openmed.service.app import OpenMedService
from openmed.processing.outputs import OutputFormatter

service = OpenMedService()
result = service.process_text("Patient was prescribed Aspirin for headache.")

formatter = OutputFormatter()
csv_rows = formatter.to_csv_rows(result)   # direct call on the result object

# Save or stream csv_rows as needed

Summary

  • OpenMed CSV customization centers on the OutputFormatter class in openmed/processing/outputs.py.
  • The to_csv_rows method converts PredictionResult objects into flat dictionaries suitable for CSV serialization.
  • Use format_predictions with output_format="csv" for quick one-off conversions, or instantiate OutputFormatter directly for pipeline integration.
  • Custom metadata automatically appears in CSV output with the metadata_ prefix.
  • The CSV format includes entity text, labels, confidence scores, character spans, model metadata, and timestamps.

Frequently Asked Questions

What fields are included in OpenMed CSV output?

The to_csv_rows method in openmed/processing/outputs.py generates dictionaries containing: text (the entity text), label (the entity type), confidence (the prediction score), start and end (character positions), model_name (the source model), timestamp (processing time), processing_time (duration), and original_text (the full input). Custom metadata fields appear with metadata_ prefixes.

How do I filter low-confidence predictions before CSV export?

Pass the confidence_threshold parameter to format_predictions or OutputFormatter.format_predictions. Values below this threshold (default 0.5) are excluded from the PredictionResult.entities list before CSV conversion occurs, ensuring only high-quality predictions appear in your output file.

Can I customize the column names in the OpenMed CSV output?

OpenMed generates standardized field names in to_csv_rows. To use custom column names, transform the returned dictionaries before writing to CSV. Iterate through csv_rows and remap keys using a dictionary comprehension, or use csv.DictWriter with the extrasaction="ignore" parameter and specify your preferred fieldnames order.

Where is the CSV formatting logic implemented in OpenMed?

The core logic resides in openmed/processing/outputs.py at lines 515+ (the to_csv_rows method) and lines 551+ (the format_predictions wrapper). The OutputFormatter class handles the conversion pipeline, while EntityPrediction (lines 55+) and PredictionResult (lines 77+) define the data structures being converted.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →