# How to Configure Sentence Detection with pySBD in OpenMED

> Configure sentence detection in OpenMED with pySBD. Learn to use language, clean, and segmenter parameters for precise medical text segmentation and improved performance.

- Repository: [Maziyar Panahi/openmed](https://github.com/maziyarpanahi/openmed)
- Tags: how-to-guide
- Published: 2026-06-11

---

**OpenMED exposes the `segment_text` function in [`openmed/processing/sentences.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/sentences.py) to split medical text into sentences using pySBD, offering three configuration parameters (`language`, `clean`, and `segmenter`) and internal caching for performance.**

OpenMED relies on the **pySBD** (Python Sentence Boundary Detector) library to segment raw clinical text into individual sentences while preserving exact character offsets. Configuring sentence detection with pySBD allows you to customize language-specific rules, preprocessing behavior, and segmenter reuse across multiple invocations. The implementation centers on a cached segmenter architecture that optimizes performance for high-volume medical text processing pipelines.

## Core Configuration Parameters

The `segment_text` function in [`openmed/processing/sentences.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/sentences.py) accepts three key parameters that control how pySBD processes your text:

| Parameter | Type | Description | Default |
|-----------|------|-------------|---------|
| `language` | `str` | ISO-639-1 language code (e.g., `"en"` for English, `"es"` for Spanish) that determines which pySBD rules to apply | `"en"` |
| `clean` | `bool` | When `True`, enables aggressive preprocessing to remove extraneous whitespace and normalize punctuation, improving accuracy on noisy clinical notes | `False` |
| `segmenter` | `Segmenter` | Optional pre-instantiated pySBD `Segmenter` object; supplying this bypasses the internal cache and allows custom configuration reuse | `None` |

When you call `segment_text` without providing a `segmenter` argument, OpenMED automatically creates and caches a segmenter based on the `language` and `clean` combination.

## Internal Caching Mechanism

OpenMED implements a **segmenter cache** (`_SEGMENTER_CACHE`) to avoid the overhead of repeatedly instantiating pySBD objects. This dictionary is keyed by tuples of `(language, clean)` and stored at the module level in [`openmed/processing/sentences.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/sentences.py).

The internal `_get_segmenter` function handles the cache logic:

1. Returns the supplied `segmenter` if one is provided.
2. Looks up the cache for a matching `(language, clean)` key and returns the cached instance.
3. If no match exists, imports `pysbd.Segmenter`, initializes it with `char_span=True` to preserve character offsets, stores it in `_SEGMENTER_CACHE`, and returns it.

```python

# From openmed/processing/sentences.py

segmenter = Segmenter(
    language=language,
    clean=clean,
    char_span=True,
)
_SEGMENTER_CACHE[cache_key] = segmenter

```

This caching strategy ensures that repeated calls with the same parameters reuse the same underlying pySBD instance, significantly reducing latency in batch processing workflows.

## Usage Examples

### Simple Sentence Segmentation

For basic use cases, let OpenMED handle segmenter creation and caching automatically:

```python
from openmed.processing.sentences import segment_text

text = "Patient is stable. No fever, no cough."
spans = segment_text(text)  # Uses default English, clean=False

for span in spans:
    print(f"Text: {span.text}, Start: {span.start}, End: {span.end}")

```

### Language and Cleaning Configuration

Process Spanish clinical notes with preprocessing enabled:

```python
spans = segment_text(
    text,
    language="es",  # Spanish language rules

    clean=True,     # Aggressive whitespace and punctuation normalization

)

```

### Custom Segmenter Reuse

For high-throughput scenarios, instantiate a custom segmenter once and reuse it across many calls to bypass cache lookups:

```python
from pysbd import Segmenter
from openmed.processing.sentences import segment_text

custom_segmenter = Segmenter(language="en", clean=False, char_span=True)

# Pass the same instance to multiple calls

first_spans = segment_text(text_batch_1, segmenter=custom_segmenter)
second_spans = segment_text(text_batch_2, segmenter=custom_segmenter)

```

### Handling Missing Dependencies

OpenMED raises a clear `ImportError` if pySBD is not installed:

```python
try:
    spans = segment_text(text)
except ImportError as exc:
    raise RuntimeError(
        "Sentence detection requires pySBD. Install with `pip install pysbd`."
    ) from exc

```

## Character Span Handling

After segmentation, OpenMED extracts both the sentence text and its **character span** (start and end offsets). When the pySBD output already contains `start` and `end` attributes, these values are used directly. Otherwise, the internal `_fallback_spans` routine constructs spans by searching the original text to ensure every sentence maintains accurate positional metadata.

This span preservation is critical for downstream clinical NLP tasks that require mapping processed sentences back to their original document positions. The integration tests in [`tests/integration/test_sentence_detection_real.py`](https://github.com/maziyarpanahi/openmed/blob/main/tests/integration/test_sentence_detection_real.py) demonstrate this behavior on real medical text samples.

## Summary

- **OpenMED** uses pySBD for sentence boundary detection via the `segment_text` function in [`openmed/processing/sentences.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/sentences.py)
- Configure detection using three parameters: `language` (ISO-639-1 code), `clean` (preprocessing toggle), and `segmenter` (custom instance)
- The module maintains a `_SEGMENTER_CACHE` keyed by `(language, clean)` tuples to optimize performance across repeated calls
- Pass `char_span=True` to pySBD to preserve character offsets, enabling accurate mapping back to original text positions
- Supply a pre-instantiated `Segmenter` to bypass caching and maintain custom configurations across batch processing

## Frequently Asked Questions

### What is the default language setting for sentence detection in OpenMED?

The default language is `"en"` (English). When you call `segment_text` without specifying the `language` parameter, OpenMED initializes a pySBD segmenter with English-specific rules. To process other languages, supply the appropriate ISO-639-1 code (e.g., `"es"` for Spanish or `"de"` for German) according to the pySBD language support.

### How does the `clean` parameter affect sentence segmentation?

When `clean=True`, pySBD performs additional preprocessing steps including removal of extraneous whitespace and normalization of punctuation marks. This mode improves segmentation accuracy on noisy clinical notes that may contain irregular spacing or formatting artifacts, though it adds slight computational overhead compared to the default `clean=False` setting.

### Can I reuse a pySBD segmenter across multiple function calls?

Yes. You can instantiate a `pysbd.Segmenter` object with your desired configuration and pass it to the `segmenter` parameter. This approach bypasses OpenMED's internal cache lookup and allows you to maintain custom settings or share a single segmenter instance across multiple threads or processes in high-throughput environments.

### What happens if pySBD is not installed in my environment?

OpenMED raises an `ImportError` with a descriptive message indicating that pySBD is required for sentence detection. You must install the dependency using `pip install pysbd` before importing the sentence processing module. The error handling is implemented in [`openmed/processing/sentences.py`](https://github.com/maziyarpanahi/openmed/blob/main/openmed/processing/sentences.py) to provide clear guidance on resolving the missing dependency.