How to Customize OpenMed Output to Dictionary Format

To customize OpenMed output to dictionary format, use the format_predictions function from openmed.processing with output_format="dict" or call .to_dict() on a PredictionResult object.

OpenMed (maziyarpanahi/openmed) ships with a flexible output-formatting subsystem that transforms raw NER model predictions into structured Python dictionaries. When you need to customize OpenMed output to dictionary format for downstream API processing or data pipelines, the library exposes dedicated classes in openmed/processing/outputs.py that handle token normalization, span correction, and entity grouping.

Understanding the OpenMed Output Architecture

The formatting pipeline centers on three core components defined in openmed/processing/outputs.py. Understanding these classes helps you configure the exact dictionary structure your application requires.

Core Dataclasses

  • EntityPrediction (lines 55-64): Holds individual entity data including text, label, confidence score, and character span offsets.
  • PredictionResult (lines 77-86): Aggregates the complete inference result, containing the original text, a list of EntityPrediction objects, model metadata, and timestamps.

The OutputFormatter Engine

The OutputFormatter class (lines 95-107) implements the heavy-lifting logic:

  1. Normalizes token text via _normalize_token_text to remove tokenizer artifacts like "▁" or "##".
  2. Corrects span offsets via _fix_entity_spans (lines 40-48) to extend short ends and trim whitespace.
  3. Groups adjacent entities via _group_adjacent_entities (lines 303-361) when the same label spans multiple tokens.
  4. Renders output to dict, json, html, or csv formats based on configuration.

The formatter is deliberately stateless; the _current_text attribute is cleared after each call (lines 19-21), preventing cross-request leakage in batch processing.

Converting Predictions to Dictionary Format

OpenMed provides two primary approaches for dictionary conversion: a convenience wrapper for quick conversions and the formatter class for reusable configurations.

Method 1: Using the format_predictions Convenience Wrapper

The format_predictions function (lines 52-93) in openmed/processing/outputs.py provides the fastest path to a dictionary:

from openmed.processing import format_predictions

raw_predictions = [
    {"entity": "B-PER", "score": 0.98, "start": 0, "end": 5, "word": "Alice"},
    {"entity": "B-LOC", "score": 0.92, "start": 10, "end": 16, "word": "London"},
]

text = "Alice lives in London."

result = format_predictions(
    predictions=raw_predictions,
    original_text=text,
    model_name="my-ner-model",
    output_format="dict",          # Request dictionary output

    include_confidence=True,       # Include scores in the dict

    confidence_threshold=0.5,      # Filter low-confidence predictions

)

# Access the dictionary

output_dict = result.to_dict()
print(output_dict)

When output_format="dict" is specified (lines 84-85), the function returns a PredictionResult instance. Calling .to_dict() on this instance (line 88) produces the final Python dictionary matching the JSON structure.

Method 2: Instantiating OutputFormatter Directly

For high-throughput applications, instantiate OutputFormatter directly to reuse configuration:

from openmed.processing.outputs import OutputFormatter

formatter = OutputFormatter(
    include_confidence=False,      # Exclude scores from output

    confidence_threshold=0.7,      # Stricter filtering

    group_entities=True            # Merge adjacent same-label entities

)

result = formatter.format_predictions(
    predictions=raw_predictions,
    original_text=text,
    model_name="my-ner-model",
    processing_time=0.023,       # Optional metadata

)

prediction_dict = result.to_dict()

This approach avoids the overhead of re-instantiating the formatter for every batch. The group_entities flag is particularly useful when tokenizers split single entities into multiple sub-tokens, as the _group_adjacent_entities method merges these fragments automatically.

Advanced Customization Options

Fine-tuning the dictionary output involves controlling how entities are filtered, grouped, and normalized before serialization.

Filtering by Confidence Threshold

The confidence_threshold parameter filters predictions during the initial processing stage. Predictions with scores below this threshold are excluded from the final PredictionResult object, reducing noise in the dictionary output.

Grouping Adjacent Entities

Set group_entities=True to enable the merging logic in _group_adjacent_entities (lines 303-361). This combines consecutive EntityPrediction objects sharing the same label into a single dictionary entry, solving the common tokenizer issue where "New York" becomes separate "New" and "York" tokens.

Token Normalization and Span Correction

Before dictionary conversion, the formatter automatically:

  • Strips tokenizer prefixes (e.g., "▁" from SentencePiece or "##" from WordPiece) via _normalize_token_text.
  • Fixes entity spans via _fix_entity_spans to ensure start/end indices align with the original text whitespace.

These corrections ensure the dictionary output contains accurate character offsets for highlighting or extraction tasks.

Complete Code Examples

Basic Dictionary Output

from openmed.processing import format_predictions

raw_predictions = [
    {"entity": "B-PER", "score": 0.98, "start": 0, "end": 5, "word": "Alice"},
    {"entity": "B-LOC", "score": 0.92, "start": 10, "end": 16, "word": "London"},
]

text = "Alice lives in London."

result = format_predictions(
    predictions=raw_predictions,
    original_text=text,
    model_name="my-ner-model",
    output_format="dict",
    include_confidence=True,
    confidence_threshold=0.5,
)

print(result.to_dict())

Custom Formatter Configuration

from openmed.processing.outputs import OutputFormatter

formatter = OutputFormatter(
    include_confidence=False,
    confidence_threshold=0.7,
    group_entities=True
)

result = formatter.format_predictions(
    predictions=raw_predictions,
    original_text=text,
    model_name="my-ner-model",
    processing_time=0.023,
)

prediction_dict = result.to_dict()
print(prediction_dict)

JSON Serialization for API Responses

json_payload = format_predictions(
    predictions=raw_predictions,
    original_text=text,
    model_name="my-ner-model",
    output_format="json",
    include_confidence=True,
)

# json_payload is a string ready for HTTP transmission

The JSON branch (lines 86-87) internally calls json.dumps on the dictionary representation produced by to_dict().

Summary

  • Use format_predictions from openmed.processing with output_format="dict" for quick conversions.
  • Call .to_dict() on PredictionResult objects to obtain the final Python dictionary.
  • Configure OutputFormatter directly when processing batches to reuse settings and improve performance.
  • Enable group_entities=True to merge adjacent tokens of the same label into single dictionary entries.
  • Set confidence_threshold to filter low-quality predictions before dictionary serialization.

Frequently Asked Questions

How do I remove confidence scores from the OpenMed dictionary output?

Pass include_confidence=False to either format_predictions or the OutputFormatter constructor. This excludes the score field from all EntityPrediction objects in the resulting dictionary, producing a cleaner output when scores are not required downstream.

Can I merge adjacent entities of the same type in OpenMed?

Yes. Set group_entities=True when calling format_predictions or initializing OutputFormatter. This triggers the _group_adjacent_entities method (lines 303-361 in outputs.py), which combines consecutive tokens with identical labels into single dictionary entries with corrected span offsets.

What is the difference between PredictionResult and EntityPrediction?

EntityPrediction (lines 55-64) represents a single extracted entity with text, label, confidence, and span coordinates. PredictionResult (lines 77-86) aggregates the entire inference output, containing the original text, a list of EntityPrediction objects, model metadata, and processing timestamps. The .to_dict() method on PredictionResult serializes the entire structure to a nested dictionary.

How do I validate entity spans after converting to dictionary format?

OpenMed provides validate_entity_spans in openmed/core/quality_gates.py (lines 202-206), which is automatically called after span fixing in the formatter. You can manually import and run this function on your dictionary output to verify that all start/end indices align with the original text boundaries and contain no overlapping or invalid spans.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →