How to Use group_entities to Merge Adjacent Entities in OpenMed
OpenMed's group_entities functionality, implemented via the _group_adjacent_entities method in openmed/processing/outputs.py, automatically consolidates fragmented entity predictions—such as dates split into "01", "/15", and "/1970"—into single cohesive entities by detecting shared labels, sentence boundaries, and proximity within two characters.
When processing PII with the OpenMed library, raw model predictions often fragment continuous entities like dates or social security numbers into separate tokens. The group_entities pipeline solves this by merging adjacent fragments before downstream processing, ensuring complete semantic units are captured as unified spans.
How _group_adjacent_entities Detects and Merges Fragments
Adjacency Detection Logic
In openmed/processing/outputs.py, the _group_adjacent_entities method iterates through EntityPrediction objects sorted by start offset and identifies adjacent pairs using three criteria implemented at lines 31-38:
- Same label: Both entities must share identical entity type labels (e.g., "date")
- Same sentence: Both entities must belong to the same sentence, verified via
metadata["sentence_index"] - Proximity: The gap between entities must be at most two characters (
entity.start <= last_entity.end + 2)
Entity Merging Process
When adjacent entities satisfy the criteria, the method accumulates them into current_group and flushes the group to _merge_entities (called at lines 42-45). The _merge_entities method performs the actual consolidation at lines 86-90:
- Stitches text spans together into a single string
- Expands character boundaries to include surrounding alphanumerics
- Calculates the arithmetic mean of confidence scores across the group
- Returns a single
EntityPredictionrepresenting the complete merged span
Practical Usage Examples
Basic Grouping of Fragmented Entities
To manually group adjacent fragments, instantiate PredictionOutput and call the internal _group_adjacent_entities method:
from openmed.processing.outputs import PredictionOutput
from openmed.core.pii import EntityPrediction
# Simulating fragmented date predictions for "DOB: 01/15/1970"
raw_entities = [
EntityPrediction(text="01", label="date", confidence=0.80, start=5, end=7,
metadata={"sentence_index": 0}),
EntityPrediction(text="/15/1970", label="date", confidence=0.75, start=7, end=15,
metadata={"sentence_index": 0}),
]
output = PredictionOutput()
grouped = output._group_adjacent_entities(raw_entities)
print(grouped[0].text) # "01/15/1970"
print(grouped[0].confidence) # Averaged score ~0.775
Full Pipeline with Semantic Merging
Combine basic grouping with semantic unit merging for context-aware pattern matching:
from openmed.core.pii import predict_pii
from openmed.core.pii_entity_merger import merge_entities_with_semantic_units
from openmed.processing.outputs import PredictionOutput
text = "Patient SSN: 123-45-6789"
# Step 1: Generate raw predictions
raw = predict_pii(text)
# Step 2: Group adjacent fragments
output = PredictionOutput()
grouped = output._group_adjacent_entities(raw)
# Step 3: Apply semantic merging (exported in openmed/__init__.py)
merged = merge_entities_with_semantic_units(
[e.__dict__ for e in grouped],
text,
use_semantic_patterns=True,
)
Disabling Automatic Grouping
To process raw token predictions without grouping, bypass PredictionOutput and feed entities directly to the semantic merger:
# Skip _group_adjacent_entities, pass raw predictions directly
merged = merge_entities_with_semantic_units(
[e.__dict__ for e in raw],
text,
use_semantic_patterns=True,
)
Key Source Files and Methods
The group_entities pipeline relies on specific implementations across the OpenMed repository:
openmed/processing/outputs.py: Contains_group_adjacent_entities(adjacency detection at lines 31-38) and_merge_entities(boundary expansion and confidence averaging at lines 86-90)openmed/core/pii_entity_merger.py: Providesmerge_entities_with_semantic_unitsfor regex-based semantic merging after basic groupingopenmed/core/pii.py: High-level PII detection entry point that automatically invokes grouping viaPredictionOutput.to_json()orto_html()
Summary
- Adjacency Requirements: Entities must share the same label and sentence index, with a gap of ≤2 characters to qualify for merging according to the logic in
openmed/processing/outputs.py. - Confidence Calculation: Merged entities receive an arithmetic mean of confidence scores from all constituent fragments.
- Automatic Invocation:
PredictionOutputmethodsto_json()andto_html()automatically call_group_adjacent_entitiesduring serialization. - Semantic Enhancement: Combine basic grouping with
merge_entities_with_semantic_unitsfor context-aware pattern matching that respects semantic boundaries. - Direct Access: While
_group_adjacent_entitiesis a private method, you can invoke it directly onPredictionOutputinstances for custom preprocessing workflows.
Frequently Asked Questions
What is the maximum character gap for merging adjacent entities?
OpenMed's adjacency detection allows a maximum gap of two characters between the end of one entity and the start of the next, as implemented in the conditional check at lines 31-38 of openmed/processing/outputs.py. This threshold captures standard tokenization fragments while preventing unrelated entities from merging.
How does OpenMed calculate confidence scores for merged entities?
The _merge_entities method calculates the arithmetic mean of confidence scores across all constituent fragments. When multiple adjacent predictions are consolidated into a single EntityPrediction, the resulting confidence score represents the average of the individual probabilities.
Can I disable automatic entity grouping in OpenMed?
Yes. While PredictionOutput.to_json() and to_html() automatically invoke _group_adjacent_entities, you can bypass grouping entirely by skipping the PredictionOutput class and feeding raw predictions directly to merge_entities_with_semantic_units or your custom post-processor.
What is the difference between _group_adjacent_entities and merge_entities_with_semantic_units?
_group_adjacent_entities performs proximity-based merging of same-label entities within the same sentence based on character offsets, while merge_entities_with_semantic_units applies regex patterns and semantic rules for context-aware merging. The latter can be used independently or as a second stage after basic grouping to combine model predictions with dictionary-based patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →