# How the GPT-SoVITS Text Preprocessing Pipeline Normalizes Chinese, Japanese, English, Korean, and Cantonese

> Discover how the GPT-SoVITS text preprocessing pipeline normalizes Chinese, Japanese, English, Korean, and Cantonese text. Learn about its modular approach for accurate phonetic conversion and BERT feature extraction.

- Repository: [RVC-Boss/GPT-SoVITS](https://github.com/RVC-Boss/GPT-SoVITS)
- Tags: internals
- Published: 2026-03-07

---

**TLDR:** The GPT-SoVITS repository implements a modular text preprocessing pipeline where the `TextPreprocessor` class dispatches input to language-specific modules—such as PaddleSpeech for Chinese, `pyopenjtalk` for Japanese, and custom IPA converters for Korean—that normalize punctuation, convert numerals, and transform text into phonetic representations before BERT feature extraction.

The RVC-Boss/GPT-SoVITS text preprocessing pipeline serves as the critical first stage in the TTS inference chain, transforming raw multilingual input into normalized phoneme sequences and acoustic feature embeddings. Before synthesizing speech, the pipeline handles the unique orthographic characteristics of Chinese, Japanese, English, Korean, and Cantonese through specialized normalization routines implemented across dedicated language modules in `GPT_SoVITS/text/`.

## Architecture of the TextPreprocessor Class

The orchestration of multilingual text normalization begins in [`GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py). The `TextPreprocessor` class coordinates segmentation, language identification, and feature extraction through a structured workflow that ensures each language follows its own normalization rules.

### Initial Segmentation and Cleaning

The `preprocess()` method acts as the primary entry point for incoming text. First, it invokes `replace_consecutive_punctuation()` to collapse redundant punctuation marks that could confuse downstream phoneme converters. Then, `pre_seg_text()` segments the input according to the specified split method (such as by sentence or natural boundary), creating discrete processing chunks that preserve linguistic coherence.

### Language Routing via LangSegmenter

For each text segment, `get_phones_and_bert()` utilizes `LangSegmenter` to detect language boundaries and assign appropriate language tags (`zh`, `ja`, `en`, `ko`, `yue`). This routing mechanism calls `clean_text_inf()`, which delegates to `clean_text()` in [`GPT_SoVITS/text/cleaner.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/text/cleaner.py). The cleaner dynamically selects the appropriate language module from a `language_module_map`, ensuring that Chinese text never passes through Japanese normalization logic and vice versa.

## Language-Specific Normalization Strategies

Each supported language implements its own `text_normalize()` function within dedicated modules under `GPT_SoVITS/text/`, following a consistent pattern of punctuation standardization, character filtering, and orthographic normalization.

### Chinese and Cantonese Text Normalization

Both Chinese (`zh`) and Cantonese (`yue`) modules rely on PaddleSpeech's `TextNormalizer` to handle numeric expansion and text segmentation. In [`text/chinese.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/chinese.py) and [`text/cantonese.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/cantonese.py), the `text_normalize()` function converts Arabic numerals to Chinese characters and standardizes punctuation through `replace_punctuation()`, which maps full-width symbols to ASCII equivalents.

Cantonese specifically utilizes the `ToJyutping` library after normalization to convert characters into Jyutping romanization before phoneme extraction. This allows the pipeline to handle Cantonese-specific pronunciations while sharing the same BERT feature extraction infrastructure as Mandarin.

### Japanese Text Cleaning

The Japanese module ([`text/japanese.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/japanese.py)) implements a lightweight normalization approach. Its `text_normalize()` function primarily deduplicates consecutive punctuation marks, while `replace_consecutive_punctuation()` ensures that multiple exclamation points or question marks collapse into single instances.

Unlike Chinese, Japanese normalization does not perform extensive numeric conversion, delegating prosodic and phonetic complexity handling directly to the `pyopenjtalk` grapheme-to-phoneme engine.

### English Normalization Pipeline

English processing in [`text/english.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/english.py) employs a regex-based `rep_map` to normalize diverse punctuation variants and special symbols. The `text_normalize()` function first converts Chinese numerals to Arabic using `cn2an`, then applies expansions from `en_normalization` to handle ordinals and cardinals.

After punctuation collapsing via `replace_consecutive_punctuation()`, the cleaned text feeds into `g2p_en` for phoneme generation.

### Korean Text Processing

The Korean module ([`text/korean.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/korean.py)) implements a unique three-stage normalization. First, `latin_to_hangul()` transcribes embedded Latin characters into Hangul. Then, `number_to_hangul()` converts Arabic numerals into spoken Korean using appropriate classifiers.

Finally, `korean_to_ipa()` utilizes `g2pk2` and `ko_pron` to generate IPA representations, which `post_replace_ph()` then maps to the repository's lazy-IPA symbol set compatible with the acoustic model.

## Grapheme-to-Phoneme Conversion

After text normalization, each language module executes its own `g2p()` function to generate phoneme lists and `word2ph` alignment maps.

- **Chinese** variants use `pypinyin` (or the optional G2PW BERT model in [`chinese2.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/chinese2.py)) to extract initials and finals before tone-sandhi application.
- **Japanese** delegates entirely to `pyopenjtalk` for mora-level phoneme extraction.
- **English** relies on `g2p_en` with custom symbol filtering via `replace_phs`.
- **Korean** maps through IPA to lazy-IPA symbols.
- **Cantonese** expands Jyutping syllables into initial-final-tone tuples prefixed with "Y" for vowel-consonant merging.

## BERT Feature Extraction

For Chinese and Cantonese text, the pipeline extracts per-phoneme BERT embeddings via `get_bert_feature()` in [`TextPreprocessor.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/TextPreprocessor.py). This function tokenizes the normalized text, runs inference through a masked language model (typically `bert-base-chinese`), and repeats each token's hidden state according to the `word2ph` alignment, producing a `[1024 × phoneme]` tensor.

Non-Chinese languages receive zero-tensor placeholders, as the acoustic model does not require BERT features for these languages.

## Practical Implementation Examples

```python
from GPT_SoVITS.TTS_infer_pack.TextPreprocessor import TextPreprocessor
from transformers import AutoModelForMaskedLM, AutoTokenizer
import torch

# Initialize BERT for Chinese/Cantonese feature extraction

bert = AutoModelForMaskedLM.from_pretrained("bert-base-chinese")
tokenizer = AutoTokenizer.from_pretrained("bert-base-chinese")
device = torch.device("cpu")

processor = TextPreprocessor(bert, tokenizer, device)

# Process mixed-language input

mixed_text = "你好，this is a test。今日は良い天気です！"
segments = processor.preprocess(mixed_text, lang="auto", text_split_method="sentence")

for seg in segments:
    print(f"Phones: {seg['phones']}")
    print(f"Normalized: {seg['norm_text']}")
    print(f"BERT shape: {seg['bert_features'].shape}")

```

```python

# Korean numeral normalization example

korean_input = "오늘은 2023년 5월 10일입니다."
result = processor.preprocess(korean_input, lang="ko", text_split_method="sentence")
print(result[0]["norm_text"])

# Output: "오늘은 이천이십삼년 오월 십일입니다."

```

## Summary

- The `TextPreprocessor` class in [`GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py) orchestrates multilingual text normalization by routing input through language-specific modules defined in [`GPT_SoVITS/text/cleaner.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/GPT_SoVITS/text/cleaner.py).
- **Chinese** and **Cantonese** utilize PaddleSpeech's `TextNormalizer` for numeric expansion and punctuation standardization, with Cantonese additionally converting characters to Jyutping phonemes via `ToJyutping`.
- **Japanese** normalization focuses on punctuation deduplication before delegating phonetic processing to `pyopenjtalk`.
- **English** employs regex-based symbol mapping and `cn2an` for numeral conversion, feeding into `g2p_en` for phoneme generation.
- **Korean** implements multi-stage normalization converting Latin characters to Hangul and numbers to spoken form, then maps through IPA to lazy-IPA symbols via `g2pk2` and `ko_pron`.
- **BERT feature extraction** occurs only for Chinese and Cantonese, generating `[1024 × phoneme]` embeddings while other languages use zero placeholders.

## Frequently Asked Questions

### How does GPT-SoVITS handle mixed-language input?

The pipeline uses `LangSegmenter` within `get_phones_and_bert()` to automatically detect language boundaries and split mixed input into homogeneous segments. Each segment then routes through its respective language module ([`chinese.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/chinese.py), [`english.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/english.py), etc.) for normalization and g2p conversion, allowing seamless processing of sentences containing multiple languages.

### What distinguishes the Chinese and Chinese2 modes?

Both modes utilize identical PaddleSpeech normalization in [`text/chinese.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/chinese.py) and [`text/chinese2.py`](https://github.com/RVC-Boss/GPT-SoVITS/blob/main/text/chinese2.py), but **Chinese2** optionally enables the G2PW BERT model for enhanced pinyin prediction when the `is_g2pw` flag is activated. This provides more accurate polyphone disambiguation compared to the standard `pypinyin` implementation used in the base Chinese mode.

### Why does Cantonese use the same BERT model as Mandarin Chinese?

Cantonese (`yue`) shares the `bert-base-chinese` model with Mandarin because the pipeline utilizes the same Chinese character-based BERT tokenizer and model architecture for both languages. The Jyutping conversion occurs after BERT feature extraction, allowing the acoustic model to leverage shared contextual embeddings despite differing phonetic realizations.

### How does the pipeline handle Arabic numerals in Korean text?

The Korean module's `number_to_hangul()` function converts Arabic numerals to their spoken Hangul equivalents using Korean number classifiers (e.g., 2023 becomes 이천이십삼). This occurs during the `text_normalize()` phase before the `g2p` stage, ensuring the TTS model receives phonetically spelled-out numbers rather than digit sequences.