How the GPT-SoVITS Text Preprocessing Pipeline Normalizes Chinese, Japanese, English, Korean, and Cantonese

TLDR: The GPT-SoVITS repository implements a modular text preprocessing pipeline where the TextPreprocessor class dispatches input to language-specific modules—such as PaddleSpeech for Chinese, pyopenjtalk for Japanese, and custom IPA converters for Korean—that normalize punctuation, convert numerals, and transform text into phonetic representations before BERT feature extraction.

The RVC-Boss/GPT-SoVITS text preprocessing pipeline serves as the critical first stage in the TTS inference chain, transforming raw multilingual input into normalized phoneme sequences and acoustic feature embeddings. Before synthesizing speech, the pipeline handles the unique orthographic characteristics of Chinese, Japanese, English, Korean, and Cantonese through specialized normalization routines implemented across dedicated language modules in GPT_SoVITS/text/.

Architecture of the TextPreprocessor Class

The orchestration of multilingual text normalization begins in GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py. The TextPreprocessor class coordinates segmentation, language identification, and feature extraction through a structured workflow that ensures each language follows its own normalization rules.

Initial Segmentation and Cleaning

The preprocess() method acts as the primary entry point for incoming text. First, it invokes replace_consecutive_punctuation() to collapse redundant punctuation marks that could confuse downstream phoneme converters. Then, pre_seg_text() segments the input according to the specified split method (such as by sentence or natural boundary), creating discrete processing chunks that preserve linguistic coherence.

Language Routing via LangSegmenter

For each text segment, get_phones_and_bert() utilizes LangSegmenter to detect language boundaries and assign appropriate language tags (zh, ja, en, ko, yue). This routing mechanism calls clean_text_inf(), which delegates to clean_text() in GPT_SoVITS/text/cleaner.py. The cleaner dynamically selects the appropriate language module from a language_module_map, ensuring that Chinese text never passes through Japanese normalization logic and vice versa.

Language-Specific Normalization Strategies

Each supported language implements its own text_normalize() function within dedicated modules under GPT_SoVITS/text/, following a consistent pattern of punctuation standardization, character filtering, and orthographic normalization.

Chinese and Cantonese Text Normalization

Both Chinese (zh) and Cantonese (yue) modules rely on PaddleSpeech's TextNormalizer to handle numeric expansion and text segmentation. In text/chinese.py and text/cantonese.py, the text_normalize() function converts Arabic numerals to Chinese characters and standardizes punctuation through replace_punctuation(), which maps full-width symbols to ASCII equivalents.

Cantonese specifically utilizes the ToJyutping library after normalization to convert characters into Jyutping romanization before phoneme extraction. This allows the pipeline to handle Cantonese-specific pronunciations while sharing the same BERT feature extraction infrastructure as Mandarin.

Japanese Text Cleaning

The Japanese module (text/japanese.py) implements a lightweight normalization approach. Its text_normalize() function primarily deduplicates consecutive punctuation marks, while replace_consecutive_punctuation() ensures that multiple exclamation points or question marks collapse into single instances.

Unlike Chinese, Japanese normalization does not perform extensive numeric conversion, delegating prosodic and phonetic complexity handling directly to the pyopenjtalk grapheme-to-phoneme engine.

English Normalization Pipeline

English processing in text/english.py employs a regex-based rep_map to normalize diverse punctuation variants and special symbols. The text_normalize() function first converts Chinese numerals to Arabic using cn2an, then applies expansions from en_normalization to handle ordinals and cardinals.

After punctuation collapsing via replace_consecutive_punctuation(), the cleaned text feeds into g2p_en for phoneme generation.

Korean Text Processing

The Korean module (text/korean.py) implements a unique three-stage normalization. First, latin_to_hangul() transcribes embedded Latin characters into Hangul. Then, number_to_hangul() converts Arabic numerals into spoken Korean using appropriate classifiers.

Finally, korean_to_ipa() utilizes g2pk2 and ko_pron to generate IPA representations, which post_replace_ph() then maps to the repository's lazy-IPA symbol set compatible with the acoustic model.

Grapheme-to-Phoneme Conversion

After text normalization, each language module executes its own g2p() function to generate phoneme lists and word2ph alignment maps.

  • Chinese variants use pypinyin (or the optional G2PW BERT model in chinese2.py) to extract initials and finals before tone-sandhi application.
  • Japanese delegates entirely to pyopenjtalk for mora-level phoneme extraction.
  • English relies on g2p_en with custom symbol filtering via replace_phs.
  • Korean maps through IPA to lazy-IPA symbols.
  • Cantonese expands Jyutping syllables into initial-final-tone tuples prefixed with "Y" for vowel-consonant merging.

BERT Feature Extraction

For Chinese and Cantonese text, the pipeline extracts per-phoneme BERT embeddings via get_bert_feature() in TextPreprocessor.py. This function tokenizes the normalized text, runs inference through a masked language model (typically bert-base-chinese), and repeats each token's hidden state according to the word2ph alignment, producing a [1024 × phoneme] tensor.

Non-Chinese languages receive zero-tensor placeholders, as the acoustic model does not require BERT features for these languages.

Practical Implementation Examples

from GPT_SoVITS.TTS_infer_pack.TextPreprocessor import TextPreprocessor
from transformers import AutoModelForMaskedLM, AutoTokenizer
import torch

# Initialize BERT for Chinese/Cantonese feature extraction

bert = AutoModelForMaskedLM.from_pretrained("bert-base-chinese")
tokenizer = AutoTokenizer.from_pretrained("bert-base-chinese")
device = torch.device("cpu")

processor = TextPreprocessor(bert, tokenizer, device)

# Process mixed-language input

mixed_text = "你好,this is a test。今日は良い天気です!"
segments = processor.preprocess(mixed_text, lang="auto", text_split_method="sentence")

for seg in segments:
    print(f"Phones: {seg['phones']}")
    print(f"Normalized: {seg['norm_text']}")
    print(f"BERT shape: {seg['bert_features'].shape}")

# Korean numeral normalization example

korean_input = "오늘은 2023년 5월 10일입니다."
result = processor.preprocess(korean_input, lang="ko", text_split_method="sentence")
print(result[0]["norm_text"])

# Output: "오늘은 이천이십삼년 오월 십일입니다."

Summary

  • The TextPreprocessor class in GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py orchestrates multilingual text normalization by routing input through language-specific modules defined in GPT_SoVITS/text/cleaner.py.
  • Chinese and Cantonese utilize PaddleSpeech's TextNormalizer for numeric expansion and punctuation standardization, with Cantonese additionally converting characters to Jyutping phonemes via ToJyutping.
  • Japanese normalization focuses on punctuation deduplication before delegating phonetic processing to pyopenjtalk.
  • English employs regex-based symbol mapping and cn2an for numeral conversion, feeding into g2p_en for phoneme generation.
  • Korean implements multi-stage normalization converting Latin characters to Hangul and numbers to spoken form, then maps through IPA to lazy-IPA symbols via g2pk2 and ko_pron.
  • BERT feature extraction occurs only for Chinese and Cantonese, generating [1024 × phoneme] embeddings while other languages use zero placeholders.

Frequently Asked Questions

How does GPT-SoVITS handle mixed-language input?

The pipeline uses LangSegmenter within get_phones_and_bert() to automatically detect language boundaries and split mixed input into homogeneous segments. Each segment then routes through its respective language module (chinese.py, english.py, etc.) for normalization and g2p conversion, allowing seamless processing of sentences containing multiple languages.

What distinguishes the Chinese and Chinese2 modes?

Both modes utilize identical PaddleSpeech normalization in text/chinese.py and text/chinese2.py, but Chinese2 optionally enables the G2PW BERT model for enhanced pinyin prediction when the is_g2pw flag is activated. This provides more accurate polyphone disambiguation compared to the standard pypinyin implementation used in the base Chinese mode.

Why does Cantonese use the same BERT model as Mandarin Chinese?

Cantonese (yue) shares the bert-base-chinese model with Mandarin because the pipeline utilizes the same Chinese character-based BERT tokenizer and model architecture for both languages. The Jyutping conversion occurs after BERT feature extraction, allowing the acoustic model to leverage shared contextual embeddings despite differing phonetic realizations.

How does the pipeline handle Arabic numerals in Korean text?

The Korean module's number_to_hangul() function converts Arabic numerals to their spoken Hangul equivalents using Korean number classifiers (e.g., 2023 becomes 이천이십삼). This occurs during the text_normalize() phase before the g2p stage, ensuring the TTS model receives phonetically spelled-out numbers rather than digit sequences.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →