Customizing OpenMed Output to JSON Format: Methods and Examples

OpenMed generates JSON output through the OutputFormatter class in openmed/processing/outputs.py, which serializes PredictionResult objects using json.dumps with configurable indentation, and you can customize this output by subclassing the formatter or extending the result's to_dict() method.

OpenMed's medical text processing pipeline produces structured predictions that require flexible serialization for downstream integration. Customizing OpenMed output to JSON format enables seamless interoperability with web APIs, data warehouses, and analytics pipelines. The library provides a robust formatting architecture centered on the OutputFormatter class, allowing both programmatic and command-line customization of JSON serialization.

Understanding the JSON Output Architecture

The JSON serialization chain in OpenMed relies on two core components working in tandem: the OutputFormatter class and the PredictionResult data structure.

The OutputFormatter Class

Located in openmed/processing/outputs.py, the OutputFormatter class serves as the primary gateway for converting prediction results into JSON strings. The to_json() method implements a thin wrapper around Python's json.dumps, accepting a PredictionResult object and an optional indentation parameter that defaults to 2 spaces.

The method delegates the initial conversion to the result's to_dict() helper, which transforms the structured data into a Python dictionary before JSON serialization. This architecture separates data representation from formatting concerns, making it straightforward to inject custom processing steps.

PredictionResult Structure

The PredictionResult class encapsulates all inference outputs, including the original text, model name, processing timestamp, optional processing time, and an array of EntityPrediction objects. Each entity contains text, label, confidence, and optional start and end character offsets. Before serialization, internal methods like _fix_entity_spans normalize these offsets to ensure consistency in the final JSON output.

Built-in JSON Serialization Methods

OpenMed provides multiple pathways to generate JSON output, accommodating both programmatic workflows and CLI usage.

Programmatic Usage with to_json()

The simplest approach involves instantiating OutputFormatter and calling to_json() with your prediction result:

from openmed.processing.outputs import OutputFormatter, PredictionResult, EntityPrediction

# Build a dummy result

result = PredictionResult(
    text="Paciente con diabetes tipo 2.",
    entities=[
        EntityPrediction(
            text="diabetes tipo 2",
            label="DISEASE",
            confidence=0.96,
            start=13,
            end=30,
        )
    ],
    model_name="disease_detection",
    timestamp="2025-10-17T00:00:00",
    processing_time=0.123,
)

formatter = OutputFormatter()
json_output = formatter.to_json(result, indent=2)
print(json_output)

CLI JSON Output

The CLI entry point in openmed/cli/main.py exposes a --output-format json flag that routes predictions through the same OutputFormatter.to_json() routine:

openmed analyze \
  --text "Paciente con hipertensión y dolor torácico." \
  --model disease_detection \
  --output-format json

HTTP API JSON Responses

The FastAPI endpoint in openmed/service/api.py handling POST /analyze utilizes the same OutputFormatter.to_json() method to return JSON responses to HTTP clients. This ensures that API consumers receive identically structured data as the CLI output, with model metadata defined in openmed/core/models.py.

Customizing OpenMed JSON Output

While the default serialization covers standard use cases, production deployments often require custom field injection, key renaming, or entity filtering.

Extending PredictionResult.to_dict()

You can modify the dictionary representation before JSON conversion by extending the PredictionResult class and overriding to_dict():

from openmed.processing.outputs import PredictionResult

class ExtendedPredictionResult(PredictionResult):
    def to_dict(self):
        data = super().to_dict()
        data["custom_metadata"] = {"source": "hospital_api", "version": "1.2"}
        return data

This approach ensures custom fields propagate through the standard OutputFormatter pipeline without requiring changes to the formatter itself.

Subclassing OutputFormatter

For more advanced customization—such as filtering entities or restructuring the output schema—subclass OutputFormatter and override to_json():

from openmed.processing.outputs import OutputFormatter

class MyJSONFormatter(OutputFormatter):
    def to_json(self, result, indent=2):
        # Convert to dict first

        data = result.to_dict()
        # Add a custom field

        data["custom_note"] = "Generated by MyJSONFormatter"
        # Use the same JSON dump logic

        return super().to_json(result, indent=indent)

formatter = MyJSONFormatter()
print(formatter.to_json(result))

Entity Processing and JSON Schema

The default JSON schema includes normalized entity offsets processed by the internal _fix_entity_spans routine in openmed/processing/outputs.py. This normalization ensures that start and end character positions align with the original text boundaries after tokenization artifacts are removed.

When customizing output, preserve this normalization by calling the parent class's entity processing methods or explicitly invoking _fix_entity_spans on your entity predictions before serialization.

Summary

  • OpenMed's JSON output is generated by the OutputFormatter class in openmed/processing/outputs.py, which wraps json.dumps around the PredictionResult.to_dict() method.
  • Default indentation is 2 spaces, configurable via the indent parameter in to_json().
  • CLI integration uses the --output-format json flag in openmed/cli/main.py to route output through the same formatter.
  • Customization strategies include extending PredictionResult.to_dict() for field additions or subclassing OutputFormatter for structural changes and entity filtering.
  • Entity offsets are normalized by _fix_entity_spans before serialization to ensure accurate character positioning.

Frequently Asked Questions

How do I change the indentation in OpenMed JSON output?

Pass the indent parameter to OutputFormatter.to_json(). The default value is 2 spaces, but you can specify any integer or None for compact output. For CLI usage, modify the formatter instantiation in your custom subclass or post-process the output string.

Can I add custom fields to the OpenMed JSON output?

Yes. Extend PredictionResult and override the to_dict() method to inject additional fields into the dictionary before serialization. Alternatively, subclass OutputFormatter and manipulate the dictionary returned by result.to_dict() before calling json.dumps.

What entity fields are included in the default JSON format?

Each entity object includes text (the extracted span), label (the prediction category), confidence (the probability score), and optional start and end character offsets. These offsets are normalized by the internal _fix_entity_spans routine in openmed/processing/outputs.py.

How does the CLI handle JSON output formatting?

The CLI entry point in openmed/cli/main.py parses the --output-format json flag and instantiates OutputFormatter to process predictions through the to_json() method. This ensures consistency between programmatic and command-line JSON output generation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →