Voice Attributes Controlled by Text Prompts in VoxCPM: Complete Guide
VoxCPM enables zero-shot control over gender, age, tone, emotion, pace, and speaking style by embedding descriptive text instructions in parentheses before the input text, eliminating the need for reference audio clips.
VoxCPM is an open-source controllable text-to-speech model developed by OpenBMB that interprets natural language descriptions to shape synthetic voice characteristics. Understanding what voice attributes can be controlled by text prompts in VoxCPM allows developers to craft precise voice personas using simple text descriptors rather than acoustic references.
How Text Prompt Control Works
VoxCPM implements voice-design conditioning by parsing free-form text descriptors wrapped in leading parentheses. When you provide a control instruction like (warm female voice), the model treats this parenthesized prefix as a conditioning signal that flows through the acoustic generation pipeline (LocEnc → TSLM → RALM → LocDiT).
The text processing occurs in src/voxcpm/cli.py through the build_final_text() function:
def build_final_text(text: str, control: str | None) -> str:
control = (control or "").strip()
return f"({control}){text}" if control else text
Source: src/voxcpm/cli.py, lines 71-73.
This function concatenates the control instruction inside parentheses with the target text, creating a format the model's encoder interprets as a latent conditioning token sequence.
Controllable Voice Attributes
According to the project documentation in README.md (line 44) and CLI implementation, VoxCPM recognizes seven distinct categories of voice attributes through natural language prompts.
Gender and Age
Control the speaker's demographic profile using descriptors for gender and age:
- Gender: "female voice", "male voice", "young woman", "masculine voice"
- Age: "young", "old", "teenage", "middle-aged", "elderly", "child"
Example: (young male voice) or (elderly female voice).
Tone and Pitch Characteristics
Modify the acoustic quality and vocal register using tone and pitch descriptors:
- Warmth: "warm", "bright", "soft", "gentle"
- Register: "deep", "high-pitched", "low-pitched", "bright"
- Timbre: "nasal", "husky", "smooth", "clear"
Example: (warm female voice) or (deep, husky voice).
Emotional Expression
Shape the affective content through emotion keywords:
- "happy", "sad", "excited", "calm"
- "angry", "nervous", "cheerful", "melancholic"
- "smiling", "serious", "playful"
Example: (cheerful tone) or (calm, gentle voice).
Speech Pace and Speed
Control the temporal rhythm using pace descriptors:
- "slightly faster", "fast-paced"
- "slow", "very slow", "measured"
- "energetic", "relaxed pace"
Example: (slightly faster, cheerful tone) as documented in README.md lines 166-168.
Stylistic and Delivery Modifiers
Fine-tune the speaking manner through style attributes:
- "warm and gentle", "dramatic", "soft"
- "sweet", "smiling", "professional"
- Any free-form phrase conveying vocal quality (e.g., "a slightly husky voice with a playful lilt")
Example: (young female voice, warm and gentle, slightly smiling) as shown in README.md line 215.
Implementation Reference
CLI Argument Handling
The command-line interface processes the --control argument in src/voxcpm/cli.py (lines 355-364), converting user input into the parenthesized format before inference. The CLI usage follows this pattern:
voxcpm design \
--text "Hello world, welcome to VoxCPM!" \
--control "young female voice, warm and gentle, slightly smiling" \
--output out.wav
Python API Integration
The core model class in src/voxcpm/model/voxcpm.py exposes the generate() method which accepts a control_instruction parameter. This parameter feeds directly into the voice-conditioning mechanism without requiring manual string formatting.
Practical Code Examples
Python API Usage
Generate speech with controlled attributes using the model's Python interface:
from voxcpm import VoxCPM
import soundfile as sf
model = VoxCPM.from_pretrained("OpenBMB/VoxCPM2", load_denoiser=False)
# Combine multiple attributes: tone, gender, pace, and emotion
text = "Hello, this is a controllable clone."
control = "warm female voice, slightly faster, cheerful tone"
wav = model.generate(text=text, control_instruction=control)
sf.write("controllable.wav", wav, model.tts_model.sample_rate)
Command Line Interface
Rapid prototyping through the CLI allows quick iteration on voice designs:
voxcpm design \
--text "Welcome to the future of speech synthesis." \
--control "deep male voice, professional tone, slightly slower" \
--output professional_voice.wav
Gradio Web Interface
For interactive voice design, implement a web UI that forwards text prompts to the model:
import gradio as gr
from voxcpm import VoxCPM
model = VoxCPM.from_pretrained("OpenBMB/VoxCPM2")
def synthesize(text, control):
wav = model.generate(text=text, control_instruction=control)
return (48000, wav)
iface = gr.Interface(
fn=synthesize,
inputs=[
gr.Textbox(label="Text"),
gr.Textbox(label="Control Instruction (e.g., 'warm female voice, cheerful')"),
],
outputs=gr.Audio(label="Generated Speech")
)
iface.launch()
Summary
- Gender and age control requires phrases like "young male" or "elderly female" according to
README.mdline 44. - Tone and pitch modifications use descriptors such as "warm", "deep", or "bright".
- Emotional states including "happy", "sad", and "excited" steer the affective qualities of synthesis.
- Pace adjustments like "slightly faster" or "slow" alter the temporal rhythm without changing content.
- Stylistic cues such as "smiling" or "nasal" provide fine-grained control over voice texture.
- The control instruction is processed through
build_final_text()insrc/voxcpm/cli.pyand injected into the model via thecontrol_instructionparameter insrc/voxcpm/model/voxcpm.py.
Frequently Asked Questions
What is the exact syntax for controlling voice attributes in VoxCPM?
Wrap your descriptive voice attributes inside parentheses preceding the input text. The CLI utility in src/voxcpm/cli.py automatically formats this via the build_final_text() function, which inserts the control string between parentheses immediately before the content text. For example: (warm female voice)Hello world.
Can I combine multiple voice attributes in a single text prompt?
Yes. VoxCPM supports comma-separated combinations of attributes within the same parenthesized instruction. As demonstrated in the CLI examples from README.md line 215, you can specify (young female voice, warm and gentle, slightly smiling, slightly faster) to simultaneously control age, gender, tone, expression, and pace.
Do I need reference audio clips to use text prompt voice control?
No. VoxCPM supports zero-shot voice design, meaning the text prompts alone can generate voices with specified attributes without requiring reference audio files. This distinguishes the control mechanism from voice cloning features that require audio samples.
Where does the control instruction connect to the model architecture?
The control instruction flows from the CLI/API layer into the model's encoder as a conditioning signal. According to the source implementation, the parenthesized text is treated as a voice-design conditioning token sequence that influences the downstream pipeline components (LocEnc → TSLM → RALM → LocDiT), ultimately steering the acoustic generation without modifying the textual content itself.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →