What Chinese Dialects Does VoxCPM2 Support? A Complete Regional Speech Guide
VoxCPM2 supports nine Chinese dialects including Sichuanese, Cantonese, Wu, Northeastern Mandarin, Henan, Shaanxi, Shandong, Tianjin, and Southern Min, using a tokenizer-free architecture that processes raw Unicode characters without dialect-specific tokenizers.
The OpenBMB/VoxCPM repository delivers a multilingual text-to-speech system designed to handle regional linguistic diversity across China. Understanding exactly which Chinese dialects VoxCPM2 supports enables developers to build localized voice applications using a single unified model. This guide explores the supported dialects, the tokenization-free implementation, and practical code examples for synthesis.
The Nine Chinese Dialects Supported by VoxCPM2
According to the project documentation in README.md (lines 55‑57), VoxCPM2 explicitly supports the following Chinese regional varieties:
- 四川话 (Sichuanese) – Southwestern Mandarin spoken in Sichuan province and Chongqing
- 粤语 (Cantonese) – Yue Chinese spoken in Guangdong, Hong Kong, and Macau
- 吴语 (Wu) – Spoken in Shanghai, Zhejiang, and parts of Jiangsu, including Shanghainese
- 东北话 (Northeastern Mandarin) – Northern Mandarin varieties spoken in Dongbei provinces
- 河南话 (Henan Mandarin) – Central Plains Mandarin spoken in Henan province
- 陕西话 (Shaanxi Mandarin) – Northwestern Mandarin including Xi'an dialect
- 山东话 (Shandong Mandarin) – Jilu Mandarin spoken throughout Shandong province
- 天津话 (Tianjin Mandarin) – Northern Mandarin with distinct tonal patterns from Tianjin
- 闽南话 (Southern Min) – Min Nan varieties spoken in Fujian and Taiwan
The model was trained on over 2 million hours of multilingual speech data encompassing all nine dialects, allowing the diffusion-autoregressive decoder to learn dialect-specific acoustic patterns without explicit dialect flags.
How VoxCPM2 Handles Chinese Dialects Without Tokenizers
Unlike traditional TTS systems requiring language-specific token vocabularies, VoxCPM2 implements a tokenizer-free architecture that directly maps raw Unicode characters to acoustic embeddings via the MiniCPM-4-based encoder.
In src/voxcpm/utils/text_normalize.py, the split_paragraph utility (lines 58‑66) automatically detects any input containing Chinese characters and assigns lang="zh". This design treats all Chinese varieties through a unified processing path, eliminating the need for dialect-specific preprocessing or language tags.
The model infers prosody and phonetics directly from character-level context. You can input text using dialect-specific orthography—such as traditional characters for Cantonese—or standard Mandarin characters that approximate regional pronunciation. The internal language model disambiguates the intended dialect based on distributional patterns learned during training.
Practical Implementation: Generating Speech in Chinese Dialects
Basic Sichuanese Generation
The generate() method accepts raw dialect text without additional configuration parameters:
from voxcpm import VoxCPM
import soundfile as sf
model = VoxCPM.from_pretrained("openbmb/VoxCPM2", load_denoiser=False)
# Example text in 四川话 (Sichuan dialect)
text_sichuan = "咋子哦!今天的天气真巴适。"
wav = model.generate(text=text_sichuan, cfg_value=2.0, inference_timesteps=10)
sf.write("sichuan.wav", wav, model.tts_model.sample_rate)
print("Saved Sichuan dialect demo → sichuan.wav")
Cantonese with Traditional Characters
For Cantonese synthesis, supply traditional Chinese characters directly to the same generation pipeline:
text_cantonese = "你點呀?依家天氣好好喎。"
wav = model.generate(text=text_cantonese, cfg_value=2.0, inference_timesteps=10)
sf.write("cantonese.wav", wav, model.tts_model.sample_rate)
Streaming API for Real-Time Dialect Synthesis
Use generate_streaming() for low-latency applications serving regional speech:
chunks = []
for chunk in model.generate_streaming(text=text_sichuan):
chunks.append(chunk)
import numpy as np
wav = np.concatenate(chunks)
sf.write("sichuan_stream.wav", wav, model.tts_model.sample_rate)
Voice Design with Wu Dialect Prompts
The model supports prompt-based voice design combined with dialect text in src/voxcpm/model/voxcpm2.py:
# Prompt includes a voice description; the following text is in 吴语 (Wu dialect)
prompt = "(柔和、温暖的上海口音)阿拉上海宁讲起话来蛮舒服格啦。"
wav = model.generate(text=prompt, cfg_value=2.5, inference_timesteps=12)
sf.write("wu_voice.wav", wav, model.tts_model.sample_rate)
Key Implementation Files for Chinese Dialect Support
Several critical files in the OpenBMB/VoxCPM repository enable seamless dialect processing:
README.md– Documents the nine supported Chinese dialects at lines 55‑57src/voxcpm/utils/text_normalize.py– Contains thesplit_paragraphfunction (lines 58‑66) that detects Chinese input and applies unified normalization for all dialectssrc/voxcpm/model/voxcpm2.py– Implements the tokenizer-free diffusion architecture processing raw Unicode characterssrc/voxcpm/core.py– Model-selection wrapper that instantiatesVoxCPM2Modelwhen the architecture hint contains"voxcpm2", ensuring dialect-capable inferencetests/test_cli.py– Validates CLI detection ofvoxcpm2architecture, confirming the dialect-supporting model loads correctly
Summary
- VoxCPM2 supports nine distinct Chinese dialects: Sichuanese, Cantonese, Wu, Northeastern Mandarin, Henan, Shaanxi, Shandong, Tianjin, and Southern Min
- The tokenizer-free architecture processes raw Unicode characters through a MiniCPM-4 encoder without dialect-specific tokenizers
- Language detection in
src/voxcpm/utils/text_normalize.pytreats all Chinese varieties uniformly under the"zh"label - The model was trained on over 2 million hours of multilingual data encompassing all supported regional varieties
- Developers can synthesize dialect speech using standard Python API calls without passing dialect-specific flags or configuration
Frequently Asked Questions
Does VoxCPM2 require separate models for each Chinese dialect?
No. VoxCPM2 uses a single unified model architecture that handles all nine dialects simultaneously. The diffusion-autoregressive decoder learned dialect-specific acoustic patterns during pre-training on over 2 million hours of multilingual data, eliminating the need for separate model downloads or switches.
Can I use simplified Chinese characters for Cantonese synthesis?
Yes. While traditional characters are native to Cantonese orthography, the tokenizer-free architecture in src/voxcpm/model/voxcpm2.py accepts either simplified or traditional Unicode characters. The model disambiguates pronunciation based on contextual cues and character distribution observed during training.
How does the model detect which dialect to generate?
The model does not rely on explicit dialect flags. Instead, it infers the intended variety from the raw character input and contextual patterns. The split_paragraph function in src/voxcpm/utils/text_normalize.py (lines 58‑66) routes all Chinese text through a unified processing path, and the character-level language model determines appropriate prosody based on training distribution.
Where is the dialect support documented in the source code?
The official list of supported dialects appears in README.md at lines 55‑57. Implementation details regarding Chinese language detection reside in src/voxcpm/utils/text_normalize.py, while the core synthesis capabilities are implemented in src/voxcpm/model/voxcpm2.py using the MiniCPM-4-based encoder.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →