Voice Attributes Controlled by Text Prompts in VoxCPM: Complete Guide

VoxCPM enables zero-shot control over gender, age, tone, emotion, pace, and speaking style by embedding descriptive text instructions in parentheses before the input text, eliminating the need for reference audio clips.

VoxCPM is an open-source controllable text-to-speech model developed by OpenBMB that interprets natural language descriptions to shape synthetic voice characteristics. Understanding what voice attributes can be controlled by text prompts in VoxCPM allows developers to craft precise voice personas using simple text descriptors rather than acoustic references.

How Text Prompt Control Works

VoxCPM implements voice-design conditioning by parsing free-form text descriptors wrapped in leading parentheses. When you provide a control instruction like (warm female voice), the model treats this parenthesized prefix as a conditioning signal that flows through the acoustic generation pipeline (LocEnc → TSLM → RALM → LocDiT).

The text processing occurs in src/voxcpm/cli.py through the build_final_text() function:

def build_final_text(text: str, control: str | None) -> str:
    control = (control or "").strip()
    return f"({control}){text}" if control else text

Source: src/voxcpm/cli.py, lines 71-73.

This function concatenates the control instruction inside parentheses with the target text, creating a format the model's encoder interprets as a latent conditioning token sequence.

Controllable Voice Attributes

According to the project documentation in README.md (line 44) and CLI implementation, VoxCPM recognizes seven distinct categories of voice attributes through natural language prompts.

Gender and Age

Control the speaker's demographic profile using descriptors for gender and age:

  • Gender: "female voice", "male voice", "young woman", "masculine voice"
  • Age: "young", "old", "teenage", "middle-aged", "elderly", "child"

Example: (young male voice) or (elderly female voice).

Tone and Pitch Characteristics

Modify the acoustic quality and vocal register using tone and pitch descriptors:

  • Warmth: "warm", "bright", "soft", "gentle"
  • Register: "deep", "high-pitched", "low-pitched", "bright"
  • Timbre: "nasal", "husky", "smooth", "clear"

Example: (warm female voice) or (deep, husky voice).

Emotional Expression

Shape the affective content through emotion keywords:

  • "happy", "sad", "excited", "calm"
  • "angry", "nervous", "cheerful", "melancholic"
  • "smiling", "serious", "playful"

Example: (cheerful tone) or (calm, gentle voice).

Speech Pace and Speed

Control the temporal rhythm using pace descriptors:

  • "slightly faster", "fast-paced"
  • "slow", "very slow", "measured"
  • "energetic", "relaxed pace"

Example: (slightly faster, cheerful tone) as documented in README.md lines 166-168.

Stylistic and Delivery Modifiers

Fine-tune the speaking manner through style attributes:

  • "warm and gentle", "dramatic", "soft"
  • "sweet", "smiling", "professional"
  • Any free-form phrase conveying vocal quality (e.g., "a slightly husky voice with a playful lilt")

Example: (young female voice, warm and gentle, slightly smiling) as shown in README.md line 215.

Implementation Reference

CLI Argument Handling

The command-line interface processes the --control argument in src/voxcpm/cli.py (lines 355-364), converting user input into the parenthesized format before inference. The CLI usage follows this pattern:

voxcpm design \
    --text "Hello world, welcome to VoxCPM!" \
    --control "young female voice, warm and gentle, slightly smiling" \
    --output out.wav

Python API Integration

The core model class in src/voxcpm/model/voxcpm.py exposes the generate() method which accepts a control_instruction parameter. This parameter feeds directly into the voice-conditioning mechanism without requiring manual string formatting.

Practical Code Examples

Python API Usage

Generate speech with controlled attributes using the model's Python interface:

from voxcpm import VoxCPM
import soundfile as sf

model = VoxCPM.from_pretrained("OpenBMB/VoxCPM2", load_denoiser=False)

# Combine multiple attributes: tone, gender, pace, and emotion

text = "Hello, this is a controllable clone."
control = "warm female voice, slightly faster, cheerful tone"

wav = model.generate(text=text, control_instruction=control)
sf.write("controllable.wav", wav, model.tts_model.sample_rate)

Command Line Interface

Rapid prototyping through the CLI allows quick iteration on voice designs:

voxcpm design \
    --text "Welcome to the future of speech synthesis." \
    --control "deep male voice, professional tone, slightly slower" \
    --output professional_voice.wav

Gradio Web Interface

For interactive voice design, implement a web UI that forwards text prompts to the model:

import gradio as gr
from voxcpm import VoxCPM

model = VoxCPM.from_pretrained("OpenBMB/VoxCPM2")

def synthesize(text, control):
    wav = model.generate(text=text, control_instruction=control)
    return (48000, wav)

iface = gr.Interface(
    fn=synthesize,
    inputs=[
        gr.Textbox(label="Text"),
        gr.Textbox(label="Control Instruction (e.g., 'warm female voice, cheerful')"),
    ],
    outputs=gr.Audio(label="Generated Speech")
)
iface.launch()

Summary

  • Gender and age control requires phrases like "young male" or "elderly female" according to README.md line 44.
  • Tone and pitch modifications use descriptors such as "warm", "deep", or "bright".
  • Emotional states including "happy", "sad", and "excited" steer the affective qualities of synthesis.
  • Pace adjustments like "slightly faster" or "slow" alter the temporal rhythm without changing content.
  • Stylistic cues such as "smiling" or "nasal" provide fine-grained control over voice texture.
  • The control instruction is processed through build_final_text() in src/voxcpm/cli.py and injected into the model via the control_instruction parameter in src/voxcpm/model/voxcpm.py.

Frequently Asked Questions

What is the exact syntax for controlling voice attributes in VoxCPM?

Wrap your descriptive voice attributes inside parentheses preceding the input text. The CLI utility in src/voxcpm/cli.py automatically formats this via the build_final_text() function, which inserts the control string between parentheses immediately before the content text. For example: (warm female voice)Hello world.

Can I combine multiple voice attributes in a single text prompt?

Yes. VoxCPM supports comma-separated combinations of attributes within the same parenthesized instruction. As demonstrated in the CLI examples from README.md line 215, you can specify (young female voice, warm and gentle, slightly smiling, slightly faster) to simultaneously control age, gender, tone, expression, and pace.

Do I need reference audio clips to use text prompt voice control?

No. VoxCPM supports zero-shot voice design, meaning the text prompts alone can generate voices with specified attributes without requiring reference audio files. This distinguishes the control mechanism from voice cloning features that require audio samples.

Where does the control instruction connect to the model architecture?

The control instruction flows from the CLI/API layer into the model's encoder as a conditioning signal. According to the source implementation, the parenthesized text is treated as a voice-design conditioning token sequence that influences the downstream pipeline components (LocEnc → TSLM → RALM → LocDiT), ultimately steering the acoustic generation without modifying the textual content itself.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →