How to Configure Sentence Detection with pySBD in OpenMED
OpenMED exposes the segment_text function in openmed/processing/sentences.py to split medical text into sentences using pySBD, offering three configuration parameters (language, clean, and segmenter) and internal caching for performance.
OpenMED relies on the pySBD (Python Sentence Boundary Detector) library to segment raw clinical text into individual sentences while preserving exact character offsets. Configuring sentence detection with pySBD allows you to customize language-specific rules, preprocessing behavior, and segmenter reuse across multiple invocations. The implementation centers on a cached segmenter architecture that optimizes performance for high-volume medical text processing pipelines.
Core Configuration Parameters
The segment_text function in openmed/processing/sentences.py accepts three key parameters that control how pySBD processes your text:
| Parameter | Type | Description | Default |
|---|---|---|---|
language |
str |
ISO-639-1 language code (e.g., "en" for English, "es" for Spanish) that determines which pySBD rules to apply |
"en" |
clean |
bool |
When True, enables aggressive preprocessing to remove extraneous whitespace and normalize punctuation, improving accuracy on noisy clinical notes |
False |
segmenter |
Segmenter |
Optional pre-instantiated pySBD Segmenter object; supplying this bypasses the internal cache and allows custom configuration reuse |
None |
When you call segment_text without providing a segmenter argument, OpenMED automatically creates and caches a segmenter based on the language and clean combination.
Internal Caching Mechanism
OpenMED implements a segmenter cache (_SEGMENTER_CACHE) to avoid the overhead of repeatedly instantiating pySBD objects. This dictionary is keyed by tuples of (language, clean) and stored at the module level in openmed/processing/sentences.py.
The internal _get_segmenter function handles the cache logic:
- Returns the supplied
segmenterif one is provided. - Looks up the cache for a matching
(language, clean)key and returns the cached instance. - If no match exists, imports
pysbd.Segmenter, initializes it withchar_span=Trueto preserve character offsets, stores it in_SEGMENTER_CACHE, and returns it.
# From openmed/processing/sentences.py
segmenter = Segmenter(
language=language,
clean=clean,
char_span=True,
)
_SEGMENTER_CACHE[cache_key] = segmenter
This caching strategy ensures that repeated calls with the same parameters reuse the same underlying pySBD instance, significantly reducing latency in batch processing workflows.
Usage Examples
Simple Sentence Segmentation
For basic use cases, let OpenMED handle segmenter creation and caching automatically:
from openmed.processing.sentences import segment_text
text = "Patient is stable. No fever, no cough."
spans = segment_text(text) # Uses default English, clean=False
for span in spans:
print(f"Text: {span.text}, Start: {span.start}, End: {span.end}")
Language and Cleaning Configuration
Process Spanish clinical notes with preprocessing enabled:
spans = segment_text(
text,
language="es", # Spanish language rules
clean=True, # Aggressive whitespace and punctuation normalization
)
Custom Segmenter Reuse
For high-throughput scenarios, instantiate a custom segmenter once and reuse it across many calls to bypass cache lookups:
from pysbd import Segmenter
from openmed.processing.sentences import segment_text
custom_segmenter = Segmenter(language="en", clean=False, char_span=True)
# Pass the same instance to multiple calls
first_spans = segment_text(text_batch_1, segmenter=custom_segmenter)
second_spans = segment_text(text_batch_2, segmenter=custom_segmenter)
Handling Missing Dependencies
OpenMED raises a clear ImportError if pySBD is not installed:
try:
spans = segment_text(text)
except ImportError as exc:
raise RuntimeError(
"Sentence detection requires pySBD. Install with `pip install pysbd`."
) from exc
Character Span Handling
After segmentation, OpenMED extracts both the sentence text and its character span (start and end offsets). When the pySBD output already contains start and end attributes, these values are used directly. Otherwise, the internal _fallback_spans routine constructs spans by searching the original text to ensure every sentence maintains accurate positional metadata.
This span preservation is critical for downstream clinical NLP tasks that require mapping processed sentences back to their original document positions. The integration tests in tests/integration/test_sentence_detection_real.py demonstrate this behavior on real medical text samples.
Summary
- OpenMED uses pySBD for sentence boundary detection via the
segment_textfunction inopenmed/processing/sentences.py - Configure detection using three parameters:
language(ISO-639-1 code),clean(preprocessing toggle), andsegmenter(custom instance) - The module maintains a
_SEGMENTER_CACHEkeyed by(language, clean)tuples to optimize performance across repeated calls - Pass
char_span=Trueto pySBD to preserve character offsets, enabling accurate mapping back to original text positions - Supply a pre-instantiated
Segmenterto bypass caching and maintain custom configurations across batch processing
Frequently Asked Questions
What is the default language setting for sentence detection in OpenMED?
The default language is "en" (English). When you call segment_text without specifying the language parameter, OpenMED initializes a pySBD segmenter with English-specific rules. To process other languages, supply the appropriate ISO-639-1 code (e.g., "es" for Spanish or "de" for German) according to the pySBD language support.
How does the clean parameter affect sentence segmentation?
When clean=True, pySBD performs additional preprocessing steps including removal of extraneous whitespace and normalization of punctuation marks. This mode improves segmentation accuracy on noisy clinical notes that may contain irregular spacing or formatting artifacts, though it adds slight computational overhead compared to the default clean=False setting.
Can I reuse a pySBD segmenter across multiple function calls?
Yes. You can instantiate a pysbd.Segmenter object with your desired configuration and pass it to the segmenter parameter. This approach bypasses OpenMED's internal cache lookup and allows you to maintain custom settings or share a single segmenter instance across multiple threads or processes in high-throughput environments.
What happens if pySBD is not installed in my environment?
OpenMED raises an ImportError with a descriptive message indicating that pySBD is required for sentence detection. You must install the dependency using pip install pysbd before importing the sentence processing module. The error handling is implemented in openmed/processing/sentences.py to provide clear guidance on resolving the missing dependency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →