# How to Use group_entities to Merge Adjacent Entities in OpenMed

> Learn how to use OpenMed's group_entities to automatically merge fragmented entity predictions. Consolidate split entities like dates into single cohesive units efficiently.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**OpenMed's `group_entities` functionality, implemented via the `_group_adjacent_entities` method in [`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py), automatically consolidates fragmented entity predictions—such as dates split into "01", "/15", and "/1970"—into single cohesive entities by detecting shared labels, sentence boundaries, and proximity within two characters.**

When processing PII with the OpenMed library, raw model predictions often fragment continuous entities like dates or social security numbers into separate tokens. The `group_entities` pipeline solves this by merging adjacent fragments before downstream processing, ensuring complete semantic units are captured as unified spans.

## How `_group_adjacent_entities` Detects and Merges Fragments

### Adjacency Detection Logic

In [`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py), the `_group_adjacent_entities` method iterates through `EntityPrediction` objects sorted by start offset and identifies adjacent pairs using three criteria implemented at lines 31-38:

- **Same label**: Both entities must share identical entity type labels (e.g., "date")
- **Same sentence**: Both entities must belong to the same sentence, verified via `metadata["sentence_index"]`
- **Proximity**: The gap between entities must be at most two characters (`entity.start <= last_entity.end + 2`)

### Entity Merging Process

When adjacent entities satisfy the criteria, the method accumulates them into `current_group` and flushes the group to `_merge_entities` (called at lines 42-45). The `_merge_entities` method performs the actual consolidation at lines 86-90:

- Stitches text spans together into a single string
- Expands character boundaries to include surrounding alphanumerics
- Calculates the arithmetic mean of confidence scores across the group
- Returns a single `EntityPrediction` representing the complete merged span

## Practical Usage Examples

### Basic Grouping of Fragmented Entities

To manually group adjacent fragments, instantiate `PredictionOutput` and call the internal `_group_adjacent_entities` method:

```python
from openmed.processing.outputs import PredictionOutput
from openmed.core.pii import EntityPrediction

# Simulating fragmented date predictions for "DOB: 01/15/1970"

raw_entities = [
    EntityPrediction(text="01", label="date", confidence=0.80, start=5, end=7,
                     metadata={"sentence_index": 0}),
    EntityPrediction(text="/15/1970", label="date", confidence=0.75, start=7, end=15,
                     metadata={"sentence_index": 0}),
]

output = PredictionOutput()
grouped = output._group_adjacent_entities(raw_entities)

print(grouped[0].text)        # "01/15/1970"

print(grouped[0].confidence)  # Averaged score ~0.775

```

### Full Pipeline with Semantic Merging

Combine basic grouping with semantic unit merging for context-aware pattern matching:

```python
from openmed.core.pii import predict_pii
from openmed.core.pii_entity_merger import merge_entities_with_semantic_units
from openmed.processing.outputs import PredictionOutput

text = "Patient SSN: 123-45-6789"

# Step 1: Generate raw predictions

raw = predict_pii(text)

# Step 2: Group adjacent fragments

output = PredictionOutput()
grouped = output._group_adjacent_entities(raw)

# Step 3: Apply semantic merging (exported in openmed/__init__.py)

merged = merge_entities_with_semantic_units(
    [e.__dict__ for e in grouped],
    text,
    use_semantic_patterns=True,
)

```

### Disabling Automatic Grouping

To process raw token predictions without grouping, bypass `PredictionOutput` and feed entities directly to the semantic merger:

```python

# Skip _group_adjacent_entities, pass raw predictions directly

merged = merge_entities_with_semantic_units(
    [e.__dict__ for e in raw],
    text,
    use_semantic_patterns=True,
)

```

## Key Source Files and Methods

The `group_entities` pipeline relies on specific implementations across the OpenMed repository:

- **[`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py)**: Contains `_group_adjacent_entities` (adjacency detection at lines 31-38) and `_merge_entities` (boundary expansion and confidence averaging at lines 86-90)
- **[`openmed/core/pii_entity_merger.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii_entity_merger.py)**: Provides `merge_entities_with_semantic_units` for regex-based semantic merging after basic grouping
- **[`openmed/core/pii.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/core/pii.py)**: High-level PII detection entry point that automatically invokes grouping via `PredictionOutput.to_json()` or `to_html()`

## Summary

- **Adjacency Requirements**: Entities must share the same label and sentence index, with a gap of ≤2 characters to qualify for merging according to the logic in [`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py).
- **Confidence Calculation**: Merged entities receive an arithmetic mean of confidence scores from all constituent fragments.
- **Automatic Invocation**: `PredictionOutput` methods `to_json()` and `to_html()` automatically call `_group_adjacent_entities` during serialization.
- **Semantic Enhancement**: Combine basic grouping with `merge_entities_with_semantic_units` for context-aware pattern matching that respects semantic boundaries.
- **Direct Access**: While `_group_adjacent_entities` is a private method, you can invoke it directly on `PredictionOutput` instances for custom preprocessing workflows.

## Frequently Asked Questions

### What is the maximum character gap for merging adjacent entities?

OpenMed's adjacency detection allows a maximum gap of **two characters** between the end of one entity and the start of the next, as implemented in the conditional check at lines 31-38 of [`openmed/processing/outputs.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/outputs.py). This threshold captures standard tokenization fragments while preventing unrelated entities from merging.

### How does OpenMed calculate confidence scores for merged entities?

The `_merge_entities` method calculates the arithmetic mean of confidence scores across all constituent fragments. When multiple adjacent predictions are consolidated into a single `EntityPrediction`, the resulting confidence score represents the average of the individual probabilities.

### Can I disable automatic entity grouping in OpenMed?

Yes. While `PredictionOutput.to_json()` and `to_html()` automatically invoke `_group_adjacent_entities`, you can bypass grouping entirely by skipping the `PredictionOutput` class and feeding raw predictions directly to `merge_entities_with_semantic_units` or your custom post-processor.

### What is the difference between `_group_adjacent_entities` and `merge_entities_with_semantic_units`?

`_group_adjacent_entities` performs proximity-based merging of same-label entities within the same sentence based on character offsets, while `merge_entities_with_semantic_units` applies regex patterns and semantic rules for context-aware merging. The latter can be used independently or as a second stage after basic grouping to combine model predictions with dictionary-based patterns.