How to Customize OpenMed Output to Dictionary Format
To customize OpenMed output to dictionary format, use the format_predictions function from openmed.processing with output_format="dict" or call .to_dict() on a PredictionResult object.
OpenMed (maziyarpanahi/openmed) ships with a flexible output-formatting subsystem that transforms raw NER model predictions into structured Python dictionaries. When you need to customize OpenMed output to dictionary format for downstream API processing or data pipelines, the library exposes dedicated classes in openmed/processing/outputs.py that handle token normalization, span correction, and entity grouping.
Understanding the OpenMed Output Architecture
The formatting pipeline centers on three core components defined in openmed/processing/outputs.py. Understanding these classes helps you configure the exact dictionary structure your application requires.
Core Dataclasses
EntityPrediction(lines 55-64): Holds individual entity data including text, label, confidence score, and character span offsets.PredictionResult(lines 77-86): Aggregates the complete inference result, containing the original text, a list ofEntityPredictionobjects, model metadata, and timestamps.
The OutputFormatter Engine
The OutputFormatter class (lines 95-107) implements the heavy-lifting logic:
- Normalizes token text via
_normalize_token_textto remove tokenizer artifacts like"▁"or"##". - Corrects span offsets via
_fix_entity_spans(lines 40-48) to extend short ends and trim whitespace. - Groups adjacent entities via
_group_adjacent_entities(lines 303-361) when the same label spans multiple tokens. - Renders output to
dict,json,html, orcsvformats based on configuration.
The formatter is deliberately stateless; the _current_text attribute is cleared after each call (lines 19-21), preventing cross-request leakage in batch processing.
Converting Predictions to Dictionary Format
OpenMed provides two primary approaches for dictionary conversion: a convenience wrapper for quick conversions and the formatter class for reusable configurations.
Method 1: Using the format_predictions Convenience Wrapper
The format_predictions function (lines 52-93) in openmed/processing/outputs.py provides the fastest path to a dictionary:
from openmed.processing import format_predictions
raw_predictions = [
{"entity": "B-PER", "score": 0.98, "start": 0, "end": 5, "word": "Alice"},
{"entity": "B-LOC", "score": 0.92, "start": 10, "end": 16, "word": "London"},
]
text = "Alice lives in London."
result = format_predictions(
predictions=raw_predictions,
original_text=text,
model_name="my-ner-model",
output_format="dict", # Request dictionary output
include_confidence=True, # Include scores in the dict
confidence_threshold=0.5, # Filter low-confidence predictions
)
# Access the dictionary
output_dict = result.to_dict()
print(output_dict)
When output_format="dict" is specified (lines 84-85), the function returns a PredictionResult instance. Calling .to_dict() on this instance (line 88) produces the final Python dictionary matching the JSON structure.
Method 2: Instantiating OutputFormatter Directly
For high-throughput applications, instantiate OutputFormatter directly to reuse configuration:
from openmed.processing.outputs import OutputFormatter
formatter = OutputFormatter(
include_confidence=False, # Exclude scores from output
confidence_threshold=0.7, # Stricter filtering
group_entities=True # Merge adjacent same-label entities
)
result = formatter.format_predictions(
predictions=raw_predictions,
original_text=text,
model_name="my-ner-model",
processing_time=0.023, # Optional metadata
)
prediction_dict = result.to_dict()
This approach avoids the overhead of re-instantiating the formatter for every batch. The group_entities flag is particularly useful when tokenizers split single entities into multiple sub-tokens, as the _group_adjacent_entities method merges these fragments automatically.
Advanced Customization Options
Fine-tuning the dictionary output involves controlling how entities are filtered, grouped, and normalized before serialization.
Filtering by Confidence Threshold
The confidence_threshold parameter filters predictions during the initial processing stage. Predictions with scores below this threshold are excluded from the final PredictionResult object, reducing noise in the dictionary output.
Grouping Adjacent Entities
Set group_entities=True to enable the merging logic in _group_adjacent_entities (lines 303-361). This combines consecutive EntityPrediction objects sharing the same label into a single dictionary entry, solving the common tokenizer issue where "New York" becomes separate "New" and "York" tokens.
Token Normalization and Span Correction
Before dictionary conversion, the formatter automatically:
- Strips tokenizer prefixes (e.g.,
"▁"from SentencePiece or"##"from WordPiece) via_normalize_token_text. - Fixes entity spans via
_fix_entity_spansto ensure start/end indices align with the original text whitespace.
These corrections ensure the dictionary output contains accurate character offsets for highlighting or extraction tasks.
Complete Code Examples
Basic Dictionary Output
from openmed.processing import format_predictions
raw_predictions = [
{"entity": "B-PER", "score": 0.98, "start": 0, "end": 5, "word": "Alice"},
{"entity": "B-LOC", "score": 0.92, "start": 10, "end": 16, "word": "London"},
]
text = "Alice lives in London."
result = format_predictions(
predictions=raw_predictions,
original_text=text,
model_name="my-ner-model",
output_format="dict",
include_confidence=True,
confidence_threshold=0.5,
)
print(result.to_dict())
Custom Formatter Configuration
from openmed.processing.outputs import OutputFormatter
formatter = OutputFormatter(
include_confidence=False,
confidence_threshold=0.7,
group_entities=True
)
result = formatter.format_predictions(
predictions=raw_predictions,
original_text=text,
model_name="my-ner-model",
processing_time=0.023,
)
prediction_dict = result.to_dict()
print(prediction_dict)
JSON Serialization for API Responses
json_payload = format_predictions(
predictions=raw_predictions,
original_text=text,
model_name="my-ner-model",
output_format="json",
include_confidence=True,
)
# json_payload is a string ready for HTTP transmission
The JSON branch (lines 86-87) internally calls json.dumps on the dictionary representation produced by to_dict().
Summary
- Use
format_predictionsfromopenmed.processingwithoutput_format="dict"for quick conversions. - Call
.to_dict()onPredictionResultobjects to obtain the final Python dictionary. - Configure
OutputFormatterdirectly when processing batches to reuse settings and improve performance. - Enable
group_entities=Trueto merge adjacent tokens of the same label into single dictionary entries. - Set
confidence_thresholdto filter low-quality predictions before dictionary serialization.
Frequently Asked Questions
How do I remove confidence scores from the OpenMed dictionary output?
Pass include_confidence=False to either format_predictions or the OutputFormatter constructor. This excludes the score field from all EntityPrediction objects in the resulting dictionary, producing a cleaner output when scores are not required downstream.
Can I merge adjacent entities of the same type in OpenMed?
Yes. Set group_entities=True when calling format_predictions or initializing OutputFormatter. This triggers the _group_adjacent_entities method (lines 303-361 in outputs.py), which combines consecutive tokens with identical labels into single dictionary entries with corrected span offsets.
What is the difference between PredictionResult and EntityPrediction?
EntityPrediction (lines 55-64) represents a single extracted entity with text, label, confidence, and span coordinates. PredictionResult (lines 77-86) aggregates the entire inference output, containing the original text, a list of EntityPrediction objects, model metadata, and processing timestamps. The .to_dict() method on PredictionResult serializes the entire structure to a nested dictionary.
How do I validate entity spans after converting to dictionary format?
OpenMed provides validate_entity_spans in openmed/core/quality_gates.py (lines 202-206), which is automatically called after span fixing in the formatter. You can manually import and run this function on your dictionary output to verify that all start/end indices align with the original text boundaries and contain no overlapping or invalid spans.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →