How to Configure Sentence Detection with pySBD in OpenMED

OpenMED exposes the segment_text function in openmed/processing/sentences.py to split medical text into sentences using pySBD, offering three configuration parameters (language, clean, and segmenter) and internal caching for performance.

OpenMED relies on the pySBD (Python Sentence Boundary Detector) library to segment raw clinical text into individual sentences while preserving exact character offsets. Configuring sentence detection with pySBD allows you to customize language-specific rules, preprocessing behavior, and segmenter reuse across multiple invocations. The implementation centers on a cached segmenter architecture that optimizes performance for high-volume medical text processing pipelines.

Core Configuration Parameters

The segment_text function in openmed/processing/sentences.py accepts three key parameters that control how pySBD processes your text:

Parameter Type Description Default
language str ISO-639-1 language code (e.g., "en" for English, "es" for Spanish) that determines which pySBD rules to apply "en"
clean bool When True, enables aggressive preprocessing to remove extraneous whitespace and normalize punctuation, improving accuracy on noisy clinical notes False
segmenter Segmenter Optional pre-instantiated pySBD Segmenter object; supplying this bypasses the internal cache and allows custom configuration reuse None

When you call segment_text without providing a segmenter argument, OpenMED automatically creates and caches a segmenter based on the language and clean combination.

Internal Caching Mechanism

OpenMED implements a segmenter cache (_SEGMENTER_CACHE) to avoid the overhead of repeatedly instantiating pySBD objects. This dictionary is keyed by tuples of (language, clean) and stored at the module level in openmed/processing/sentences.py.

The internal _get_segmenter function handles the cache logic:

  1. Returns the supplied segmenter if one is provided.
  2. Looks up the cache for a matching (language, clean) key and returns the cached instance.
  3. If no match exists, imports pysbd.Segmenter, initializes it with char_span=True to preserve character offsets, stores it in _SEGMENTER_CACHE, and returns it.

# From openmed/processing/sentences.py

segmenter = Segmenter(
    language=language,
    clean=clean,
    char_span=True,
)
_SEGMENTER_CACHE[cache_key] = segmenter

This caching strategy ensures that repeated calls with the same parameters reuse the same underlying pySBD instance, significantly reducing latency in batch processing workflows.

Usage Examples

Simple Sentence Segmentation

For basic use cases, let OpenMED handle segmenter creation and caching automatically:

from openmed.processing.sentences import segment_text

text = "Patient is stable. No fever, no cough."
spans = segment_text(text)  # Uses default English, clean=False

for span in spans:
    print(f"Text: {span.text}, Start: {span.start}, End: {span.end}")

Language and Cleaning Configuration

Process Spanish clinical notes with preprocessing enabled:

spans = segment_text(
    text,
    language="es",  # Spanish language rules

    clean=True,     # Aggressive whitespace and punctuation normalization

)

Custom Segmenter Reuse

For high-throughput scenarios, instantiate a custom segmenter once and reuse it across many calls to bypass cache lookups:

from pysbd import Segmenter
from openmed.processing.sentences import segment_text

custom_segmenter = Segmenter(language="en", clean=False, char_span=True)

# Pass the same instance to multiple calls

first_spans = segment_text(text_batch_1, segmenter=custom_segmenter)
second_spans = segment_text(text_batch_2, segmenter=custom_segmenter)

Handling Missing Dependencies

OpenMED raises a clear ImportError if pySBD is not installed:

try:
    spans = segment_text(text)
except ImportError as exc:
    raise RuntimeError(
        "Sentence detection requires pySBD. Install with `pip install pysbd`."
    ) from exc

Character Span Handling

After segmentation, OpenMED extracts both the sentence text and its character span (start and end offsets). When the pySBD output already contains start and end attributes, these values are used directly. Otherwise, the internal _fallback_spans routine constructs spans by searching the original text to ensure every sentence maintains accurate positional metadata.

This span preservation is critical for downstream clinical NLP tasks that require mapping processed sentences back to their original document positions. The integration tests in tests/integration/test_sentence_detection_real.py demonstrate this behavior on real medical text samples.

Summary

  • OpenMED uses pySBD for sentence boundary detection via the segment_text function in openmed/processing/sentences.py
  • Configure detection using three parameters: language (ISO-639-1 code), clean (preprocessing toggle), and segmenter (custom instance)
  • The module maintains a _SEGMENTER_CACHE keyed by (language, clean) tuples to optimize performance across repeated calls
  • Pass char_span=True to pySBD to preserve character offsets, enabling accurate mapping back to original text positions
  • Supply a pre-instantiated Segmenter to bypass caching and maintain custom configurations across batch processing

Frequently Asked Questions

What is the default language setting for sentence detection in OpenMED?

The default language is "en" (English). When you call segment_text without specifying the language parameter, OpenMED initializes a pySBD segmenter with English-specific rules. To process other languages, supply the appropriate ISO-639-1 code (e.g., "es" for Spanish or "de" for German) according to the pySBD language support.

How does the clean parameter affect sentence segmentation?

When clean=True, pySBD performs additional preprocessing steps including removal of extraneous whitespace and normalization of punctuation marks. This mode improves segmentation accuracy on noisy clinical notes that may contain irregular spacing or formatting artifacts, though it adds slight computational overhead compared to the default clean=False setting.

Can I reuse a pySBD segmenter across multiple function calls?

Yes. You can instantiate a pysbd.Segmenter object with your desired configuration and pass it to the segmenter parameter. This approach bypasses OpenMED's internal cache lookup and allows you to maintain custom settings or share a single segmenter instance across multiple threads or processes in high-throughput environments.

What happens if pySBD is not installed in my environment?

OpenMED raises an ImportError with a descriptive message indicating that pySBD is required for sentence detection. You must install the dependency using pip install pysbd before importing the sentence processing module. The error handling is implemented in openmed/processing/sentences.py to provide clear guidance on resolving the missing dependency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →