How to Configure Output Formatting in OpenMed: Dict, JSON, HTML, and CSV
Use the output_format parameter in format_predictions() to return clinical entity predictions as a Python dict, JSON string, HTML fragment, or list of CSV rows.
OpenMed provides flexible output formatting for its clinical NLP pipeline, allowing you to configure exactly how prediction results are structured for downstream consumption. The format_predictions function in openmed/processing/outputs.py serves as the primary entry point for converting raw model outputs into your desired representation.
Using the format_predictions Function
The format_predictions function accepts raw token-level predictions from any Hugging Face compatible pipeline and wraps them in a structured output format. The function signature at lines 52-66 handles four distinct output types through a single configurable parameter.
from openmed import format_predictions
result = format_predictions(
raw_predictions,
original_text=text,
model_name="disease_detection_superclinical",
output_format="dict", # Options: "dict", "json", "html", "csv"
include_confidence=True,
confidence_threshold=0.2,
group_entities=False,
)
When output_format is omitted, it defaults to "dict". The function instantiates an OutputFormatter internally and delegates to OutputFormatter.format_predictions before converting to the final representation.
Output Format Options
OpenMed supports four interchangeable formats, each optimized for specific downstream workflows.
Python Dictionary (dict)
Setting output_format="dict" returns a PredictionResult dataclass instance defined at lines 77-92 in openmed/processing/outputs.py. This provides programmatic access to entity attributes and supports further serialization.
result = format_predictions(preds, txt, model_name="my_model", output_format="dict")
print(result.entities[0].label) # "DISEASE"
print(result.entities[0].text) # "chronic myeloid leukemia"
# Convert to plain dict later if needed
plain_dict = result.to_dict()
Use this format for custom aggregation logic, chaining multiple models, or when you need to access entity metadata like confidence scores directly through Python attributes.
JSON String (json)
Setting output_format="json" returns a formatted JSON string suitable for API responses or file storage. Internally, this calls OutputFormatter.to_json, which executes json.dumps(result.to_dict(), indent=2) as implemented in lines 13-23.
json_str = format_predictions(preds, txt, model_name="my_model", output_format="json")
print(json_str[:120])
# {"text":"Patient started on imatinib ...","entities":[{"text":"imatinib","label":"DRUG", ...
This format is ideal for HTTP responses, logging to NoSQL databases, or writing results to .json files using standard file operations.
HTML Fragment (html)
Setting output_format="html" generates a styled HTML string for visualization in Jupyter notebooks or web dashboards. The OutputFormatter.to_html method (lines 25-65) constructs a <div> containing the original text with inline <span> tags highlighting detected entities.
html = format_predictions(preds, txt, model_name="my_model", output_format="html")
from IPython.display import HTML
HTML(html) # Renders colored entity highlights
The HTML output includes entity-specific styling and a summary list, making it suitable for interactive dashboards and clinical review interfaces.
CSV Rows (csv)
Setting output_format="csv" returns a List[Dict[str, Any]] where each dictionary represents a row containing entity details. As implemented in lines 151-176, each row includes the entity text, label, confidence scores, and original text reference.
rows = format_predictions(preds, txt, model_name="my_model", output_format="csv")
import pandas as pd
df = pd.DataFrame(rows)
df.to_csv("entities.csv", index=False)
This format facilitates bulk export to spreadsheets, pandas analysis, or ingestion into data warehouses.
Customizing Formatter Behavior
The OutputFormatter accepts three boolean flags that control content across all formats. Pass these as keyword arguments to format_predictions:
include_confidence: Embed confidence scores in HTML summaries and CSV rowsconfidence_threshold: Filter entities below the specified probability (e.g.,0.5)group_entities: Merge adjacent entities of the same label using the internal_group_adjacent_entitieslogic
result = format_predictions(
preds,
txt,
model_name="my_model",
output_format="json",
include_confidence=False,
confidence_threshold=0.7,
group_entities=True,
processing_time=0.12, # Extra kwargs become metadata in PredictionResult
)
Any additional keyword arguments beyond these three are captured as metadata fields within the PredictionResult object.
Source Code Reference
Key implementation files in the maziyarpanahi/openmed repository:
| File | Key Symbols | Purpose |
|---|---|---|
openmed/processing/outputs.py |
PredictionResult, OutputFormatter, format_predictions |
Core formatting engine and conversion helpers |
openmed/core/quality_gates.py |
validate_entity_spans |
Post-processing validation of entity boundaries |
openmed/utils/validation.py |
validate_output_format |
Input validation ensuring only supported format strings are accepted |
Summary
- Use
output_format="dict"for programmatic access toPredictionResultobjects with full attribute access - Use
output_format="json"for API responses and file storage viajson.dumps()serialization - Use
output_format="html"for visualization in notebooks and web dashboards with inline entity highlighting - Use
output_format="csv"for pandas DataFrame conversion and spreadsheet export - Control content quality using
include_confidence,confidence_threshold, andgroup_entitiesparameters - Reference implementation lives in
openmed/processing/outputs.pylines 52-92 and 151-176
Frequently Asked Questions
How do I convert OpenMed predictions to a pandas DataFrame?
Pass output_format="csv" to format_predictions, which returns a list of dictionaries compatible with pandas. Import pandas and call pd.DataFrame(rows) to create a DataFrame, then use df.to_csv() for file export.
What is the default output format if I don't specify one?
The output_format parameter defaults to "dict", returning a PredictionResult dataclass instance. This provides the most flexibility for further processing but requires calling .to_dict() or .to_json() manually if you need serialization.
Can I filter out low-confidence predictions in the HTML output?
Yes. Pass confidence_threshold=0.5 (or your desired threshold) to format_predictions. This parameter is passed to the OutputFormatter constructor and filters entities before formatting, regardless of whether you choose HTML, JSON, dict, or CSV output.
Where does OpenMed validate the output format string?
The repository includes validation logic in openmed/utils/validation.py via the validate_output_format function, ensuring only "dict", "json", "html", or "csv" are accepted before processing begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →