How to Configure Output Formatting in OpenMed: Dict, JSON, HTML, and CSV

Use the output_format parameter in format_predictions() to return clinical entity predictions as a Python dict, JSON string, HTML fragment, or list of CSV rows.

OpenMed provides flexible output formatting for its clinical NLP pipeline, allowing you to configure exactly how prediction results are structured for downstream consumption. The format_predictions function in openmed/processing/outputs.py serves as the primary entry point for converting raw model outputs into your desired representation.

Using the format_predictions Function

The format_predictions function accepts raw token-level predictions from any Hugging Face compatible pipeline and wraps them in a structured output format. The function signature at lines 52-66 handles four distinct output types through a single configurable parameter.

from openmed import format_predictions

result = format_predictions(
    raw_predictions,
    original_text=text,
    model_name="disease_detection_superclinical",
    output_format="dict",  # Options: "dict", "json", "html", "csv"

    include_confidence=True,
    confidence_threshold=0.2,
    group_entities=False,
)

When output_format is omitted, it defaults to "dict". The function instantiates an OutputFormatter internally and delegates to OutputFormatter.format_predictions before converting to the final representation.

Output Format Options

OpenMed supports four interchangeable formats, each optimized for specific downstream workflows.

Python Dictionary (dict)

Setting output_format="dict" returns a PredictionResult dataclass instance defined at lines 77-92 in openmed/processing/outputs.py. This provides programmatic access to entity attributes and supports further serialization.

result = format_predictions(preds, txt, model_name="my_model", output_format="dict")
print(result.entities[0].label)  # "DISEASE"

print(result.entities[0].text)  # "chronic myeloid leukemia"

# Convert to plain dict later if needed

plain_dict = result.to_dict()

Use this format for custom aggregation logic, chaining multiple models, or when you need to access entity metadata like confidence scores directly through Python attributes.

JSON String (json)

Setting output_format="json" returns a formatted JSON string suitable for API responses or file storage. Internally, this calls OutputFormatter.to_json, which executes json.dumps(result.to_dict(), indent=2) as implemented in lines 13-23.

json_str = format_predictions(preds, txt, model_name="my_model", output_format="json")
print(json_str[:120])

# {"text":"Patient started on imatinib ...","entities":[{"text":"imatinib","label":"DRUG", ...

This format is ideal for HTTP responses, logging to NoSQL databases, or writing results to .json files using standard file operations.

HTML Fragment (html)

Setting output_format="html" generates a styled HTML string for visualization in Jupyter notebooks or web dashboards. The OutputFormatter.to_html method (lines 25-65) constructs a <div> containing the original text with inline <span> tags highlighting detected entities.

html = format_predictions(preds, txt, model_name="my_model", output_format="html")
from IPython.display import HTML
HTML(html)  # Renders colored entity highlights

The HTML output includes entity-specific styling and a summary list, making it suitable for interactive dashboards and clinical review interfaces.

CSV Rows (csv)

Setting output_format="csv" returns a List[Dict[str, Any]] where each dictionary represents a row containing entity details. As implemented in lines 151-176, each row includes the entity text, label, confidence scores, and original text reference.

rows = format_predictions(preds, txt, model_name="my_model", output_format="csv")
import pandas as pd
df = pd.DataFrame(rows)
df.to_csv("entities.csv", index=False)

This format facilitates bulk export to spreadsheets, pandas analysis, or ingestion into data warehouses.

Customizing Formatter Behavior

The OutputFormatter accepts three boolean flags that control content across all formats. Pass these as keyword arguments to format_predictions:

  • include_confidence: Embed confidence scores in HTML summaries and CSV rows
  • confidence_threshold: Filter entities below the specified probability (e.g., 0.5)
  • group_entities: Merge adjacent entities of the same label using the internal _group_adjacent_entities logic
result = format_predictions(
    preds,
    txt,
    model_name="my_model",
    output_format="json",
    include_confidence=False,
    confidence_threshold=0.7,
    group_entities=True,
    processing_time=0.12,  # Extra kwargs become metadata in PredictionResult

)

Any additional keyword arguments beyond these three are captured as metadata fields within the PredictionResult object.

Source Code Reference

Key implementation files in the maziyarpanahi/openmed repository:

File Key Symbols Purpose
openmed/processing/outputs.py PredictionResult, OutputFormatter, format_predictions Core formatting engine and conversion helpers
openmed/core/quality_gates.py validate_entity_spans Post-processing validation of entity boundaries
openmed/utils/validation.py validate_output_format Input validation ensuring only supported format strings are accepted

Summary

  • Use output_format="dict" for programmatic access to PredictionResult objects with full attribute access
  • Use output_format="json" for API responses and file storage via json.dumps() serialization
  • Use output_format="html" for visualization in notebooks and web dashboards with inline entity highlighting
  • Use output_format="csv" for pandas DataFrame conversion and spreadsheet export
  • Control content quality using include_confidence, confidence_threshold, and group_entities parameters
  • Reference implementation lives in openmed/processing/outputs.py lines 52-92 and 151-176

Frequently Asked Questions

How do I convert OpenMed predictions to a pandas DataFrame?

Pass output_format="csv" to format_predictions, which returns a list of dictionaries compatible with pandas. Import pandas and call pd.DataFrame(rows) to create a DataFrame, then use df.to_csv() for file export.

What is the default output format if I don't specify one?

The output_format parameter defaults to "dict", returning a PredictionResult dataclass instance. This provides the most flexibility for further processing but requires calling .to_dict() or .to_json() manually if you need serialization.

Can I filter out low-confidence predictions in the HTML output?

Yes. Pass confidence_threshold=0.5 (or your desired threshold) to format_predictions. This parameter is passed to the OutputFormatter constructor and filters entities before formatting, regardless of whether you choose HTML, JSON, dict, or CSV output.

Where does OpenMed validate the output format string?

The repository includes validation logic in openmed/utils/validation.py via the validate_output_format function, ensuring only "dict", "json", "html", or "csv" are accepted before processing begins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →