Customizing OpenMed Output to JSON Format: Methods and Examples
OpenMed generates JSON output through the OutputFormatter class in openmed/processing/outputs.py, which serializes PredictionResult objects using json.dumps with configurable indentation, and you can customize this output by subclassing the formatter or extending the result's to_dict() method.
OpenMed's medical text processing pipeline produces structured predictions that require flexible serialization for downstream integration. Customizing OpenMed output to JSON format enables seamless interoperability with web APIs, data warehouses, and analytics pipelines. The library provides a robust formatting architecture centered on the OutputFormatter class, allowing both programmatic and command-line customization of JSON serialization.
Understanding the JSON Output Architecture
The JSON serialization chain in OpenMed relies on two core components working in tandem: the OutputFormatter class and the PredictionResult data structure.
The OutputFormatter Class
Located in openmed/processing/outputs.py, the OutputFormatter class serves as the primary gateway for converting prediction results into JSON strings. The to_json() method implements a thin wrapper around Python's json.dumps, accepting a PredictionResult object and an optional indentation parameter that defaults to 2 spaces.
The method delegates the initial conversion to the result's to_dict() helper, which transforms the structured data into a Python dictionary before JSON serialization. This architecture separates data representation from formatting concerns, making it straightforward to inject custom processing steps.
PredictionResult Structure
The PredictionResult class encapsulates all inference outputs, including the original text, model name, processing timestamp, optional processing time, and an array of EntityPrediction objects. Each entity contains text, label, confidence, and optional start and end character offsets. Before serialization, internal methods like _fix_entity_spans normalize these offsets to ensure consistency in the final JSON output.
Built-in JSON Serialization Methods
OpenMed provides multiple pathways to generate JSON output, accommodating both programmatic workflows and CLI usage.
Programmatic Usage with to_json()
The simplest approach involves instantiating OutputFormatter and calling to_json() with your prediction result:
from openmed.processing.outputs import OutputFormatter, PredictionResult, EntityPrediction
# Build a dummy result
result = PredictionResult(
text="Paciente con diabetes tipo 2.",
entities=[
EntityPrediction(
text="diabetes tipo 2",
label="DISEASE",
confidence=0.96,
start=13,
end=30,
)
],
model_name="disease_detection",
timestamp="2025-10-17T00:00:00",
processing_time=0.123,
)
formatter = OutputFormatter()
json_output = formatter.to_json(result, indent=2)
print(json_output)
CLI JSON Output
The CLI entry point in openmed/cli/main.py exposes a --output-format json flag that routes predictions through the same OutputFormatter.to_json() routine:
openmed analyze \
--text "Paciente con hipertensión y dolor torácico." \
--model disease_detection \
--output-format json
HTTP API JSON Responses
The FastAPI endpoint in openmed/service/api.py handling POST /analyze utilizes the same OutputFormatter.to_json() method to return JSON responses to HTTP clients. This ensures that API consumers receive identically structured data as the CLI output, with model metadata defined in openmed/core/models.py.
Customizing OpenMed JSON Output
While the default serialization covers standard use cases, production deployments often require custom field injection, key renaming, or entity filtering.
Extending PredictionResult.to_dict()
You can modify the dictionary representation before JSON conversion by extending the PredictionResult class and overriding to_dict():
from openmed.processing.outputs import PredictionResult
class ExtendedPredictionResult(PredictionResult):
def to_dict(self):
data = super().to_dict()
data["custom_metadata"] = {"source": "hospital_api", "version": "1.2"}
return data
This approach ensures custom fields propagate through the standard OutputFormatter pipeline without requiring changes to the formatter itself.
Subclassing OutputFormatter
For more advanced customization—such as filtering entities or restructuring the output schema—subclass OutputFormatter and override to_json():
from openmed.processing.outputs import OutputFormatter
class MyJSONFormatter(OutputFormatter):
def to_json(self, result, indent=2):
# Convert to dict first
data = result.to_dict()
# Add a custom field
data["custom_note"] = "Generated by MyJSONFormatter"
# Use the same JSON dump logic
return super().to_json(result, indent=indent)
formatter = MyJSONFormatter()
print(formatter.to_json(result))
Entity Processing and JSON Schema
The default JSON schema includes normalized entity offsets processed by the internal _fix_entity_spans routine in openmed/processing/outputs.py. This normalization ensures that start and end character positions align with the original text boundaries after tokenization artifacts are removed.
When customizing output, preserve this normalization by calling the parent class's entity processing methods or explicitly invoking _fix_entity_spans on your entity predictions before serialization.
Summary
- OpenMed's JSON output is generated by the
OutputFormatterclass inopenmed/processing/outputs.py, which wrapsjson.dumpsaround thePredictionResult.to_dict()method. - Default indentation is 2 spaces, configurable via the
indentparameter into_json(). - CLI integration uses the
--output-format jsonflag inopenmed/cli/main.pyto route output through the same formatter. - Customization strategies include extending
PredictionResult.to_dict()for field additions or subclassingOutputFormatterfor structural changes and entity filtering. - Entity offsets are normalized by
_fix_entity_spansbefore serialization to ensure accurate character positioning.
Frequently Asked Questions
How do I change the indentation in OpenMed JSON output?
Pass the indent parameter to OutputFormatter.to_json(). The default value is 2 spaces, but you can specify any integer or None for compact output. For CLI usage, modify the formatter instantiation in your custom subclass or post-process the output string.
Can I add custom fields to the OpenMed JSON output?
Yes. Extend PredictionResult and override the to_dict() method to inject additional fields into the dictionary before serialization. Alternatively, subclass OutputFormatter and manipulate the dictionary returned by result.to_dict() before calling json.dumps.
What entity fields are included in the default JSON format?
Each entity object includes text (the extracted span), label (the prediction category), confidence (the probability score), and optional start and end character offsets. These offsets are normalized by the internal _fix_entity_spans routine in openmed/processing/outputs.py.
How does the CLI handle JSON output formatting?
The CLI entry point in openmed/cli/main.py parses the --output-format json flag and instantiates OutputFormatter to process predictions through the to_json() method. This ensures consistency between programmatic and command-line JSON output generation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →