# TextPreprocessor Class in GPT-SoVITS: How It Segments Text for TTS

> Explore the TextPreprocessor class in GPT-SoVITS. Learn how it segments text using language-aware techniques for efficient TTS input preparation.

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: internals
- Published: 2026-03-07

---

**The `TextPreprocessor` class transforms raw user input into phoneme-indexed tensors with contextual BERT embeddings by applying language-aware segmentation, punctuation normalization, and length-aware fragment optimization.**

The `TextPreprocessor` class serves as the primary entry point for text preparation in the GPT-SoVITS TTS pipeline. It bridges the gap between raw user sentences and the low-level representations required by the acoustic model. By orchestrating cleaning, segmentation, and feature extraction, this component ensures that text is optimally structured for high-quality speech synthesis.

## What Is the TextPreprocessor Class?

The `TextPreprocessor` class, defined in [`GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py), functions as the preprocessing engine for the GPT-SoVITS text-to-speech system. Its core responsibility is to convert unstructured text into a structured format containing phoneme IDs and BERT embeddings.

According to the RVC-Boss/GPT-SoVITS source code, the class executes four primary operations:

- **Punctuation cleaning** – Collapses consecutive punctuation marks via `replace_consecutive_punctuation` (lines 35-39)
- **Language-aware pre-segmentation** – Handles sentence boundaries and fragmentation strategies in `pre_seg_text` (lines 77-115)
- **Phone and BERT extraction** – Generates phoneme sequences and contextual embeddings through `segment_and_extract_feature_for_text` and `get_phones_and_bert` (lines 115-190)
- **Data packaging** – Returns dictionaries containing `phones`, `bert_features`, and normalized text for the acoustic model

## How TextPreprocessor Segments Text for TTS

Text segmentation follows a deterministic pipeline designed to balance linguistic coherence with model constraints. The process occurs within the `pre_seg_text` method and handles mixed-language input, punctuation normalization, and length limitations.

### Step 1: Text Cleaning and Normalization

The pipeline begins by sanitizing raw input. The `replace_consecutive_punctuation` method uses regex to collapse repeated punctuation characters into single instances (lines 35-39):

```python
punctuations = "".join(re.escape(p) for p in punctuation)
pattern = f"([{punctuations}])([{punctuations}])+"
result = re.sub(pattern, r"\1", text)

```

Initial trimming removes leading newlines (line 78), and the system ensures text begins with proper sentence-ending punctuation if the initial fragment is shorter than four characters (lines 81-82).

### Step 2: Language-Aware Pre-Segmentation

The class selects a segmentation strategy dynamically based on the `text_split_method` argument (e.g., `sentence`, `comma`, or `custom`). At line 86, `get_seg_method` resolves the appropriate function, which is then applied to the cleaned text at line 87:

```python
seg_method = get_seg_method(text_split_method)
text = seg_method(text)

```

Double newlines are collapsed into single instances (lines 89-91), and the text splits on newline delimiters to create tentative fragments (line 92).

### Step 3: Fragment Optimization

The system filters out empty lines and pure symbol fragments using `filter_text` (line 93). It then merges very short fragments (fewer than 5 characters) to prevent sub-optimal tiny utterances:

```python
_texts = merge_short_text_in_array(_texts, 5)

```

Each fragment is guaranteed to end with a sentence terminator—`。` for non-English text or `.` for English (lines 104-106). For fragments exceeding BERT's token limit of approximately 510 tokens, the `split_big_text` helper divides them into manageable pieces (lines 108-110):

```python
if len(text) > 510:
    texts.extend(split_big_text(text))
else:
    texts.append(text)

```

### Step 4: Feature Extraction Pipeline

After segmentation, the `preprocess` method iterates over the fragment list. For each segment, it calls `segment_and_extract_feature_for_text`, which invokes `get_phones_and_bert` to:

1. Run language-specific cleaners from [`GPT_SoVITS/text/cleaner.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/text/cleaner.py)
2. Convert cleaned phonemes to integer IDs via `cleaned_text_to_sequence`
3. Generate BERT embeddings with shape `(1024, num_phones)` for contextual representation

The final output is a list of dictionaries, each containing the phoneme sequence, BERT feature tensor, and normalized text string.

## Code Example: Using TextPreprocessor for TTS Inference

The following example demonstrates how to instantiate and use the class for inference:

```python
import torch
from transformers import AutoModelForMaskedLM, AutoTokenizer
from GPT_SoVITS.TTS_infer_pack.TextPreprocessor import TextPreprocessor

# Load BERT model and tokenizer

bert_model = AutoModelForMaskedLM.from_pretrained("bert-base-multilingual-cased")
tokenizer = AutoTokenizer.from_pretrained("bert-base-multilingual-cased")
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
bert_model.to(device)

# Initialize preprocessor

preproc = TextPreprocessor(bert_model, tokenizer, device)

# Process mixed-language text

raw_text = "你好，我是ChatGPT！This is a test. 这是一段很长的中文文本。"
segments = preproc.preprocess(
    text=raw_text,
    lang="auto",
    text_split_method="sentence",
    version="v2"
)

# Inspect output

print("Phones:", segments[0]["phones"])
print("BERT shape:", segments[0]["bert_features"].shape)
print("Normalized:", segments[0]["norm_text"])

```

## Key Source Files and Implementation Details

Understanding the complete pipeline requires examining these specific files in the RVC-Boss/GPT-SoVITS repository:

- **[`GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py)** – Core class implementation containing `preprocess`, `pre_seg_text`, and feature extraction logic
- **[`GPT_SoVITS/TTS_infer_pack/text_segmentation_method.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/TTS_infer_pack/text_segmentation_method.py)** – Defines segmentation strategies including sentence and comma-based splitters
- **[`GPT_SoVITS/text/LangSegmenter/LangSegmenter.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/text/LangSegmenter/LangSegmenter.py)** – Handles multilingual language detection and tagging
- **[`GPT_SoVITS/text/cleaner.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/text/cleaner.py)** – Performs language-specific text normalization before phoneme conversion
- **[`GPT_SoVITS/text/__init__.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/text/__init__.py)** – Exports `cleaned_text_to_sequence` for phoneme-to-ID mapping

## Summary

- The `TextPreprocessor` class converts raw text into model-ready phoneme tensors and BERT embeddings
- Segmentation occurs in `pre_seg_text` via punctuation normalization, method selection, and length-aware fragment optimization
- The system merges fragments shorter than 5 characters and splits fragments longer than 510 tokens to respect BERT constraints
- Language-aware processing supports multilingual input through the `lang="auto"` parameter and specialized cleaners
- Output dictionaries contain `phones` (integer IDs), `bert_features` (shape 1024×N), and `norm_text` for acoustic model consumption

## Frequently Asked Questions

### What is the maximum text length the TextPreprocessor can handle per segment?

The TextPreprocessor enforces a soft limit of approximately 510 tokens per segment to accommodate BERT's maximum sequence length. Fragments exceeding this threshold are automatically split using the `split_big_text` helper function before feature extraction occurs.

### How does TextPreprocessor handle mixed Chinese and English input?

The class utilizes the `LangSegmenter` module to auto-detect language boundaries when `lang="auto"` is specified. Each detected language segment is routed to its appropriate cleaner in [`GPT_SoVITS/text/cleaner.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/text/cleaner.py), ensuring proper phoneme conversion for both Chinese characters and English words within the same input string.

### Can I customize how TextPreprocessor splits sentences?

Yes. The `text_split_method` parameter accepts configurable strategies defined in [`text_segmentation_method.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text_segmentation_method.py). You can select `sentence` for standard sentence boundaries, `comma` for comma-based splitting, or implement custom segmentation logic by extending the segmentation method registry.

### Why does TextPreprocessor merge short text fragments?

Fragments containing fewer than 5 characters are merged with adjacent segments to prevent the TTS model from generating unnatural, choppy utterances. This optimization occurs in `merge_short_text_in_array` and improves prosodic continuity in the final synthesized speech.