How to Use group_entities to Merge Adjacent Entities in OpenMed

OpenMed's group_entities functionality, implemented via the _group_adjacent_entities method in openmed/processing/outputs.py, automatically consolidates fragmented entity predictions—such as dates split into "01", "/15", and "/1970"—into single cohesive entities by detecting shared labels, sentence boundaries, and proximity within two characters.

When processing PII with the OpenMed library, raw model predictions often fragment continuous entities like dates or social security numbers into separate tokens. The group_entities pipeline solves this by merging adjacent fragments before downstream processing, ensuring complete semantic units are captured as unified spans.

How _group_adjacent_entities Detects and Merges Fragments

Adjacency Detection Logic

In openmed/processing/outputs.py, the _group_adjacent_entities method iterates through EntityPrediction objects sorted by start offset and identifies adjacent pairs using three criteria implemented at lines 31-38:

  • Same label: Both entities must share identical entity type labels (e.g., "date")
  • Same sentence: Both entities must belong to the same sentence, verified via metadata["sentence_index"]
  • Proximity: The gap between entities must be at most two characters (entity.start <= last_entity.end + 2)

Entity Merging Process

When adjacent entities satisfy the criteria, the method accumulates them into current_group and flushes the group to _merge_entities (called at lines 42-45). The _merge_entities method performs the actual consolidation at lines 86-90:

  • Stitches text spans together into a single string
  • Expands character boundaries to include surrounding alphanumerics
  • Calculates the arithmetic mean of confidence scores across the group
  • Returns a single EntityPrediction representing the complete merged span

Practical Usage Examples

Basic Grouping of Fragmented Entities

To manually group adjacent fragments, instantiate PredictionOutput and call the internal _group_adjacent_entities method:

from openmed.processing.outputs import PredictionOutput
from openmed.core.pii import EntityPrediction

# Simulating fragmented date predictions for "DOB: 01/15/1970"

raw_entities = [
    EntityPrediction(text="01", label="date", confidence=0.80, start=5, end=7,
                     metadata={"sentence_index": 0}),
    EntityPrediction(text="/15/1970", label="date", confidence=0.75, start=7, end=15,
                     metadata={"sentence_index": 0}),
]

output = PredictionOutput()
grouped = output._group_adjacent_entities(raw_entities)

print(grouped[0].text)        # "01/15/1970"

print(grouped[0].confidence)  # Averaged score ~0.775

Full Pipeline with Semantic Merging

Combine basic grouping with semantic unit merging for context-aware pattern matching:

from openmed.core.pii import predict_pii
from openmed.core.pii_entity_merger import merge_entities_with_semantic_units
from openmed.processing.outputs import PredictionOutput

text = "Patient SSN: 123-45-6789"

# Step 1: Generate raw predictions

raw = predict_pii(text)

# Step 2: Group adjacent fragments

output = PredictionOutput()
grouped = output._group_adjacent_entities(raw)

# Step 3: Apply semantic merging (exported in openmed/__init__.py)

merged = merge_entities_with_semantic_units(
    [e.__dict__ for e in grouped],
    text,
    use_semantic_patterns=True,
)

Disabling Automatic Grouping

To process raw token predictions without grouping, bypass PredictionOutput and feed entities directly to the semantic merger:


# Skip _group_adjacent_entities, pass raw predictions directly

merged = merge_entities_with_semantic_units(
    [e.__dict__ for e in raw],
    text,
    use_semantic_patterns=True,
)

Key Source Files and Methods

The group_entities pipeline relies on specific implementations across the OpenMed repository:

  • openmed/processing/outputs.py: Contains _group_adjacent_entities (adjacency detection at lines 31-38) and _merge_entities (boundary expansion and confidence averaging at lines 86-90)
  • openmed/core/pii_entity_merger.py: Provides merge_entities_with_semantic_units for regex-based semantic merging after basic grouping
  • openmed/core/pii.py: High-level PII detection entry point that automatically invokes grouping via PredictionOutput.to_json() or to_html()

Summary

  • Adjacency Requirements: Entities must share the same label and sentence index, with a gap of ≤2 characters to qualify for merging according to the logic in openmed/processing/outputs.py.
  • Confidence Calculation: Merged entities receive an arithmetic mean of confidence scores from all constituent fragments.
  • Automatic Invocation: PredictionOutput methods to_json() and to_html() automatically call _group_adjacent_entities during serialization.
  • Semantic Enhancement: Combine basic grouping with merge_entities_with_semantic_units for context-aware pattern matching that respects semantic boundaries.
  • Direct Access: While _group_adjacent_entities is a private method, you can invoke it directly on PredictionOutput instances for custom preprocessing workflows.

Frequently Asked Questions

What is the maximum character gap for merging adjacent entities?

OpenMed's adjacency detection allows a maximum gap of two characters between the end of one entity and the start of the next, as implemented in the conditional check at lines 31-38 of openmed/processing/outputs.py. This threshold captures standard tokenization fragments while preventing unrelated entities from merging.

How does OpenMed calculate confidence scores for merged entities?

The _merge_entities method calculates the arithmetic mean of confidence scores across all constituent fragments. When multiple adjacent predictions are consolidated into a single EntityPrediction, the resulting confidence score represents the average of the individual probabilities.

Can I disable automatic entity grouping in OpenMed?

Yes. While PredictionOutput.to_json() and to_html() automatically invoke _group_adjacent_entities, you can bypass grouping entirely by skipping the PredictionOutput class and feeding raw predictions directly to merge_entities_with_semantic_units or your custom post-processor.

What is the difference between _group_adjacent_entities and merge_entities_with_semantic_units?

_group_adjacent_entities performs proximity-based merging of same-label entities within the same sentence based on character offsets, while merge_entities_with_semantic_units applies regex patterns and semantic rules for context-aware merging. The latter can be used independently or as a second stage after basic grouping to combine model predictions with dictionary-based patterns.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →