How the GPT-SoVITS Text Preprocessing Pipeline Normalizes Chinese, Japanese, English, Korean, and Cantonese
TLDR: The GPT-SoVITS repository implements a modular text preprocessing pipeline where the TextPreprocessor class dispatches input to language-specific modules—such as PaddleSpeech for Chinese, pyopenjtalk for Japanese, and custom IPA converters for Korean—that normalize punctuation, convert numerals, and transform text into phonetic representations before BERT feature extraction.
The RVC-Boss/GPT-SoVITS text preprocessing pipeline serves as the critical first stage in the TTS inference chain, transforming raw multilingual input into normalized phoneme sequences and acoustic feature embeddings. Before synthesizing speech, the pipeline handles the unique orthographic characteristics of Chinese, Japanese, English, Korean, and Cantonese through specialized normalization routines implemented across dedicated language modules in GPT_SoVITS/text/.
Architecture of the TextPreprocessor Class
The orchestration of multilingual text normalization begins in GPT_SoVITS/TTS_infer_pack/TextPreprocessor.py. The TextPreprocessor class coordinates segmentation, language identification, and feature extraction through a structured workflow that ensures each language follows its own normalization rules.
Initial Segmentation and Cleaning
The preprocess() method acts as the primary entry point for incoming text. First, it invokes replace_consecutive_punctuation() to collapse redundant punctuation marks that could confuse downstream phoneme converters. Then, pre_seg_text() segments the input according to the specified split method (such as by sentence or natural boundary), creating discrete processing chunks that preserve linguistic coherence.
Language Routing via LangSegmenter
For each text segment, get_phones_and_bert() utilizes LangSegmenter to detect language boundaries and assign appropriate language tags (zh, ja, en, ko, yue). This routing mechanism calls clean_text_inf(), which delegates to clean_text() in GPT_SoVITS/text/cleaner.py. The cleaner dynamically selects the appropriate language module from a language_module_map, ensuring that Chinese text never passes through Japanese normalization logic and vice versa.
Language-Specific Normalization Strategies
Each supported language implements its own text_normalize() function within dedicated modules under GPT_SoVITS/text/, following a consistent pattern of punctuation standardization, character filtering, and orthographic normalization.
Chinese and Cantonese Text Normalization
Both Chinese (zh) and Cantonese (yue) modules rely on PaddleSpeech's TextNormalizer to handle numeric expansion and text segmentation. In text/chinese.py and text/cantonese.py, the text_normalize() function converts Arabic numerals to Chinese characters and standardizes punctuation through replace_punctuation(), which maps full-width symbols to ASCII equivalents.
Cantonese specifically utilizes the ToJyutping library after normalization to convert characters into Jyutping romanization before phoneme extraction. This allows the pipeline to handle Cantonese-specific pronunciations while sharing the same BERT feature extraction infrastructure as Mandarin.
Japanese Text Cleaning
The Japanese module (text/japanese.py) implements a lightweight normalization approach. Its text_normalize() function primarily deduplicates consecutive punctuation marks, while replace_consecutive_punctuation() ensures that multiple exclamation points or question marks collapse into single instances.
Unlike Chinese, Japanese normalization does not perform extensive numeric conversion, delegating prosodic and phonetic complexity handling directly to the pyopenjtalk grapheme-to-phoneme engine.
English Normalization Pipeline
English processing in text/english.py employs a regex-based rep_map to normalize diverse punctuation variants and special symbols. The text_normalize() function first converts Chinese numerals to Arabic using cn2an, then applies expansions from en_normalization to handle ordinals and cardinals.
After punctuation collapsing via replace_consecutive_punctuation(), the cleaned text feeds into g2p_en for phoneme generation.
Korean Text Processing
The Korean module (text/korean.py) implements a unique three-stage normalization. First, latin_to_hangul() transcribes embedded Latin characters into Hangul. Then, number_to_hangul() converts Arabic numerals into spoken Korean using appropriate classifiers.
Finally, korean_to_ipa() utilizes g2pk2 and ko_pron to generate IPA representations, which post_replace_ph() then maps to the repository's lazy-IPA symbol set compatible with the acoustic model.
Grapheme-to-Phoneme Conversion
After text normalization, each language module executes its own g2p() function to generate phoneme lists and word2ph alignment maps.
- Chinese variants use
pypinyin(or the optional G2PW BERT model inchinese2.py) to extract initials and finals before tone-sandhi application. - Japanese delegates entirely to
pyopenjtalkfor mora-level phoneme extraction. - English relies on
g2p_enwith custom symbol filtering viareplace_phs. - Korean maps through IPA to lazy-IPA symbols.
- Cantonese expands Jyutping syllables into initial-final-tone tuples prefixed with "Y" for vowel-consonant merging.
BERT Feature Extraction
For Chinese and Cantonese text, the pipeline extracts per-phoneme BERT embeddings via get_bert_feature() in TextPreprocessor.py. This function tokenizes the normalized text, runs inference through a masked language model (typically bert-base-chinese), and repeats each token's hidden state according to the word2ph alignment, producing a [1024 × phoneme] tensor.
Non-Chinese languages receive zero-tensor placeholders, as the acoustic model does not require BERT features for these languages.
Practical Implementation Examples
from GPT_SoVITS.TTS_infer_pack.TextPreprocessor import TextPreprocessor
from transformers import AutoModelForMaskedLM, AutoTokenizer
import torch
# Initialize BERT for Chinese/Cantonese feature extraction
bert = AutoModelForMaskedLM.from_pretrained("bert-base-chinese")
tokenizer = AutoTokenizer.from_pretrained("bert-base-chinese")
device = torch.device("cpu")
processor = TextPreprocessor(bert, tokenizer, device)
# Process mixed-language input
mixed_text = "你好,this is a test。今日は良い天気です!"
segments = processor.preprocess(mixed_text, lang="auto", text_split_method="sentence")
for seg in segments:
print(f"Phones: {seg['phones']}")
print(f"Normalized: {seg['norm_text']}")
print(f"BERT shape: {seg['bert_features'].shape}")
# Korean numeral normalization example
korean_input = "오늘은 2023년 5월 10일입니다."
result = processor.preprocess(korean_input, lang="ko", text_split_method="sentence")
print(result[0]["norm_text"])
# Output: "오늘은 이천이십삼년 오월 십일입니다."
Summary
- The
TextPreprocessorclass inGPT_SoVITS/TTS_infer_pack/TextPreprocessor.pyorchestrates multilingual text normalization by routing input through language-specific modules defined inGPT_SoVITS/text/cleaner.py. - Chinese and Cantonese utilize PaddleSpeech's
TextNormalizerfor numeric expansion and punctuation standardization, with Cantonese additionally converting characters to Jyutping phonemes viaToJyutping. - Japanese normalization focuses on punctuation deduplication before delegating phonetic processing to
pyopenjtalk. - English employs regex-based symbol mapping and
cn2anfor numeral conversion, feeding intog2p_enfor phoneme generation. - Korean implements multi-stage normalization converting Latin characters to Hangul and numbers to spoken form, then maps through IPA to lazy-IPA symbols via
g2pk2andko_pron. - BERT feature extraction occurs only for Chinese and Cantonese, generating
[1024 × phoneme]embeddings while other languages use zero placeholders.
Frequently Asked Questions
How does GPT-SoVITS handle mixed-language input?
The pipeline uses LangSegmenter within get_phones_and_bert() to automatically detect language boundaries and split mixed input into homogeneous segments. Each segment then routes through its respective language module (chinese.py, english.py, etc.) for normalization and g2p conversion, allowing seamless processing of sentences containing multiple languages.
What distinguishes the Chinese and Chinese2 modes?
Both modes utilize identical PaddleSpeech normalization in text/chinese.py and text/chinese2.py, but Chinese2 optionally enables the G2PW BERT model for enhanced pinyin prediction when the is_g2pw flag is activated. This provides more accurate polyphone disambiguation compared to the standard pypinyin implementation used in the base Chinese mode.
Why does Cantonese use the same BERT model as Mandarin Chinese?
Cantonese (yue) shares the bert-base-chinese model with Mandarin because the pipeline utilizes the same Chinese character-based BERT tokenizer and model architecture for both languages. The Jyutping conversion occurs after BERT feature extraction, allowing the acoustic model to leverage shared contextual embeddings despite differing phonetic realizations.
How does the pipeline handle Arabic numerals in Korean text?
The Korean module's number_to_hangul() function converts Arabic numerals to their spoken Hangul equivalents using Korean number classifiers (e.g., 2023 becomes 이천이십삼). This occurs during the text_normalize() phase before the g2p stage, ensuring the TTS model receives phonetically spelled-out numbers rather than digit sequences.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →