# How to Configure Output Formatting in OpenMed: Dict, JSON, HTML, and CSV

> Configure OpenMed output formats like dict, JSON, HTML, and CSV using the output_format parameter in format_predictions(). Easily extract clinical entity predictions in your desired format.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**Use the `output_format` parameter in `format_predictions()` to return clinical entity predictions as a Python dict, JSON string, HTML fragment, or list of CSV rows.**

OpenMed provides flexible output formatting for its clinical NLP pipeline, allowing you to configure exactly how prediction results are structured for downstream consumption. The `format_predictions` function in [`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py) serves as the primary entry point for converting raw model outputs into your desired representation.

## Using the format_predictions Function

The `format_predictions` function accepts raw token-level predictions from any Hugging Face compatible pipeline and wraps them in a structured output format. The function signature at lines 52-66 handles four distinct output types through a single configurable parameter.

```python
from openmed import format_predictions

result = format_predictions(
    raw_predictions,
    original_text=text,
    model_name="disease_detection_superclinical",
    output_format="dict",  # Options: "dict", "json", "html", "csv"

    include_confidence=True,
    confidence_threshold=0.2,
    group_entities=False,
)

```

When `output_format` is omitted, it defaults to `"dict"`. The function instantiates an `OutputFormatter` internally and delegates to `OutputFormatter.format_predictions` before converting to the final representation.

## Output Format Options

OpenMed supports four interchangeable formats, each optimized for specific downstream workflows.

### Python Dictionary (dict)

Setting `output_format="dict"` returns a `PredictionResult` dataclass instance defined at lines 77-92 in [`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py). This provides programmatic access to entity attributes and supports further serialization.

```python
result = format_predictions(preds, txt, model_name="my_model", output_format="dict")
print(result.entities[0].label)  # "DISEASE"

print(result.entities[0].text)  # "chronic myeloid leukemia"

# Convert to plain dict later if needed

plain_dict = result.to_dict()

```

Use this format for custom aggregation logic, chaining multiple models, or when you need to access entity metadata like confidence scores directly through Python attributes.

### JSON String (json)

Setting `output_format="json"` returns a formatted JSON string suitable for API responses or file storage. Internally, this calls `OutputFormatter.to_json`, which executes `json.dumps(result.to_dict(), indent=2)` as implemented in lines 13-23.

```python
json_str = format_predictions(preds, txt, model_name="my_model", output_format="json")
print(json_str[:120])

# {"text":"Patient started on imatinib ...","entities":[{"text":"imatinib","label":"DRUG", ...

```

This format is ideal for HTTP responses, logging to NoSQL databases, or writing results to `.json` files using standard file operations.

### HTML Fragment (html)

Setting `output_format="html"` generates a styled HTML string for visualization in Jupyter notebooks or web dashboards. The `OutputFormatter.to_html` method (lines 25-65) constructs a `<div>` containing the original text with inline `<span>` tags highlighting detected entities.

```python
html = format_predictions(preds, txt, model_name="my_model", output_format="html")
from IPython.display import HTML
HTML(html)  # Renders colored entity highlights

```

The HTML output includes entity-specific styling and a summary list, making it suitable for interactive dashboards and clinical review interfaces.

### CSV Rows (csv)

Setting `output_format="csv"` returns a `List[Dict[str, Any]]` where each dictionary represents a row containing entity details. As implemented in lines 151-176, each row includes the entity text, label, confidence scores, and original text reference.

```python
rows = format_predictions(preds, txt, model_name="my_model", output_format="csv")
import pandas as pd
df = pd.DataFrame(rows)
df.to_csv("entities.csv", index=False)

```

This format facilitates bulk export to spreadsheets, pandas analysis, or ingestion into data warehouses.

## Customizing Formatter Behavior

The `OutputFormatter` accepts three boolean flags that control content across all formats. Pass these as keyword arguments to `format_predictions`:

- **`include_confidence`**: Embed confidence scores in HTML summaries and CSV rows
- **`confidence_threshold`**: Filter entities below the specified probability (e.g., `0.5`)
- **`group_entities`**: Merge adjacent entities of the same label using the internal `_group_adjacent_entities` logic

```python
result = format_predictions(
    preds,
    txt,
    model_name="my_model",
    output_format="json",
    include_confidence=False,
    confidence_threshold=0.7,
    group_entities=True,
    processing_time=0.12,  # Extra kwargs become metadata in PredictionResult

)

```

Any additional keyword arguments beyond these three are captured as metadata fields within the `PredictionResult` object.

## Source Code Reference

Key implementation files in the `maziyarpanahi/openmed` repository:

| File | Key Symbols | Purpose |
|------|-------------|---------|
| [`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py) | `PredictionResult`, `OutputFormatter`, `format_predictions` | Core formatting engine and conversion helpers |
| [`openmed/core/quality_gates.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/quality_gates.py) | `validate_entity_spans` | Post-processing validation of entity boundaries |
| [`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py) | `validate_output_format` | Input validation ensuring only supported format strings are accepted |

## Summary

- **Use `output_format="dict"`** for programmatic access to `PredictionResult` objects with full attribute access
- **Use `output_format="json"`** for API responses and file storage via `json.dumps()` serialization
- **Use `output_format="html"`** for visualization in notebooks and web dashboards with inline entity highlighting
- **Use `output_format="csv"`** for pandas DataFrame conversion and spreadsheet export
- **Control content quality** using `include_confidence`, `confidence_threshold`, and `group_entities` parameters
- **Reference implementation** lives in [`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py) lines 52-92 and 151-176

## Frequently Asked Questions

### How do I convert OpenMed predictions to a pandas DataFrame?

Pass `output_format="csv"` to `format_predictions`, which returns a list of dictionaries compatible with pandas. Import pandas and call `pd.DataFrame(rows)` to create a DataFrame, then use `df.to_csv()` for file export.

### What is the default output format if I don't specify one?

The `output_format` parameter defaults to `"dict"`, returning a `PredictionResult` dataclass instance. This provides the most flexibility for further processing but requires calling `.to_dict()` or `.to_json()` manually if you need serialization.

### Can I filter out low-confidence predictions in the HTML output?

Yes. Pass `confidence_threshold=0.5` (or your desired threshold) to `format_predictions`. This parameter is passed to the `OutputFormatter` constructor and filters entities before formatting, regardless of whether you choose HTML, JSON, dict, or CSV output.

### Where does OpenMed validate the output format string?

The repository includes validation logic in [`openmed/utils/validation.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/utils/validation.py) via the `validate_output_format` function, ensuring only `"dict"`, `"json"`, `"html"`, or `"csv"` are accepted before processing begins.