How to Create Custom Voices from Text Descriptions in VoxCPM Voice Design
VoxCPM Voice Design generates custom voices from text descriptions by prepending control instructions inside parentheses to the target text, allowing the VoxCPM2 model to synthesize novel voices without reference audio.
VoxCPM is an open-source text-to-speech framework by OpenBMB that supports voice cloning and zero-shot voice creation. Unlike traditional voice cloning that requires a reference audio sample, VoxCPM Voice Design creates entirely new voices from natural language descriptions such as "warm female voice" or "young surfer dude."
How Voice Design Works
The Voice Design feature relies on a Control Instruction mechanism. When you provide a text description of the desired voice characteristics, the system wraps this description in parentheses and concatenates it with your target text. This combined string acts as a conditioning signal for the latent diffusion model.
The key innovation is that VoxCPM2 was trained on paired control-instruction and audio examples. During inference, the model interprets the text inside the parentheses as acoustic conditioning, extrapolating from learned patterns to generate a voice that matches the description—even if that specific voice never appeared in the training data.
Implementation Architecture
CLI Parsing
In src/voxcpm/cli.py, the --control argument is processed by the build_final_text function at lines 71-73:
def build_final_text(text: str, control: str | None) -> str:
control = (control or "").strip()
return f"({control}){text}" if control else text
This function transforms a control instruction like "warm female voice" and target text "Hello world" into the formatted string "(warm female voice)Hello world".
Core Generation Pipeline
The formatted text flows into src/voxcpm/core.py, where the VoxCPM.generate method handles the synthesis. At lines 55-63, the code calls:
generate_result = self.tts_model._generate_with_prompt_cache(
target_text=text,
prompt_cache=fixed_prompt_cache,
# ...
)
When fixed_prompt_cache is None (indicating no reference audio), the model treats the entire formatted string—including the control token—as the synthesis prompt.
Model Conditioning
Inside src/voxcpm/model/voxcpm2.py, the model parses the input sequence and detects the (<any text>) pattern. This control token routes through a dedicated control encoder that influences both the latent diffusion (LocDiT) and language model streams. The conditioning shapes every generation step, determining the final timbre, pitch, and speaking style of the output audio.
Practical Usage Examples
Command-Line Interface
The fastest way to create a custom voice is using the CLI:
voxcpm design \
--text "Hello world, this is a custom voice." \
--control "warm female voice" \
--output warm_female.wav
This command invokes build_final_text internally, formats the control instruction, and executes the generation pipeline. The output warm_female.wav contains a newly synthesized voice matching your description.
Python API Integration
For programmatic control, use the Python API directly:
from voxcpm.core import VoxCPM
# Initialize the model
model = VoxCPM.from_pretrained(
hf_model_id="openbmb/VoxCPM2",
load_denoiser=False,
)
# Format the control instruction
def build_final_text(text: str, control: str | None) -> str:
control = (control or "").strip()
return f"({control}){text}" if control else text
control = "young surfer dude, relaxed, slightly nasal"
target = "Hey, catch that wave! It's a perfect day."
final_text = build_final_text(target, control)
# Generate audio
audio = model.generate(
text=final_text,
cfg_value=2.5,
inference_timesteps=12,
)
# Save output
import soundfile as sf
sf.write("surfer.wav", audio, samplerate=model.tts_model.sample_rate)
This approach gives you fine-grained control over diffusion parameters like cfg_value (guidance strength) and inference_timesteps.
Gradio Web Interface
To use the graphical interface:
- Launch the server with
python app.py - Navigate to
http://localhost:7860 - Enter your Control Instruction (e.g., "A soft, sweet teenage girl")
- Input your Target Text
- Click Generate Speech
The Gradio UI in app.py maps the control field to the same build_final_text logic used by the CLI, ensuring consistent behavior across interfaces.
Key Source Files
src/voxcpm/cli.py– Parses the--controlargument and formats instructions viabuild_final_text(lines 71-73)src/voxcpm/core.py– Contains theVoxCPMclass and generation pipeline entry point (lines 55-63)src/voxcpm/model/voxcpm2.py– Implements the model architecture that processes control tokens during diffusionapp.py– Gradio interface exposing the Voice Design workflowtests/test_cli.py– Unit tests verifying control argument embedding (lines 154-160)
Summary
- VoxCPM Voice Design creates custom voices from text descriptions without reference audio
- Control instructions use the format
(<description>)<target_text>via thebuild_final_textfunction insrc/voxcpm/cli.py - The VoxCPM2 model interprets parenthetical control tokens as conditioning signals for its latent diffusion architecture
- You can access this functionality through the CLI, Python API, or Gradio web interface
Frequently Asked Questions
What is the exact syntax for control instructions in VoxCPM?
Wrap your voice description in parentheses immediately preceding the text you want spoken. For example: "(deep male voice)Welcome to the system." The build_final_text function in src/voxcpm/cli.py handles this formatting automatically when you use the --control argument.
Can I combine voice design with voice cloning?
No, these are mutually exclusive modes. When you provide a control instruction, fixed_prompt_cache is set to None in src/voxcpm/core.py, forcing the model to generate a voice from text description only. To clone a specific speaker, you must omit the control instruction and provide reference audio instead.
How detailed should the text description be?
VoxCPM2 responds to natural language descriptions including gender, age, emotional tone, and speaking style. Descriptions like "warm female voice" or "young surfer dude, relaxed, slightly nasal" work effectively. The model extrapolates acoustic characteristics from the training data associations with these descriptive phrases.
Where is the control token processing implemented in the source code?
The control token parsing occurs in src/voxcpm/model/voxcpm2.py, where the model detects the parenthetical pattern and routes it through the control encoder. This encoder influences both the LocDiT diffusion process and the language model components to shape the final audio output.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →