# Voice Attributes Controlled by Text Prompts in VoxCPM: Complete Guide

> Discover how VoxCPM lets you control voice gender age tone emotion pace and style with simple text prompts eliminating the need for reference audio. Master voice generation today.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: how-to-guide
- Published: 2026-04-10

---

**VoxCPM enables zero-shot control over gender, age, tone, emotion, pace, and speaking style by embedding descriptive text instructions in parentheses before the input text, eliminating the need for reference audio clips.**

VoxCPM is an open-source controllable text-to-speech model developed by OpenBMB that interprets natural language descriptions to shape synthetic voice characteristics. Understanding what voice attributes can be controlled by text prompts in VoxCPM allows developers to craft precise voice personas using simple text descriptors rather than acoustic references.

## How Text Prompt Control Works

VoxCPM implements **voice-design conditioning** by parsing free-form text descriptors wrapped in leading parentheses. When you provide a control instruction like `(warm female voice)`, the model treats this parenthesized prefix as a conditioning signal that flows through the acoustic generation pipeline (LocEnc → TSLM → RALM → LocDiT).

The text processing occurs in [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py) through the `build_final_text()` function:

```python
def build_final_text(text: str, control: str | None) -> str:
    control = (control or "").strip()
    return f"({control}){text}" if control else text

```

*Source:* [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py), lines 71-73.

This function concatenates the control instruction inside parentheses with the target text, creating a format the model's encoder interprets as a latent conditioning token sequence.

## Controllable Voice Attributes

According to the project documentation in [`README.md`](https://github.com/OpenBMB/VoxCPM/blob/main/README.md) (line 44) and CLI implementation, VoxCPM recognizes seven distinct categories of voice attributes through natural language prompts.

### Gender and Age

Control the speaker's demographic profile using descriptors for **gender** and **age**:

- **Gender**: "female voice", "male voice", "young woman", "masculine voice"
- **Age**: "young", "old", "teenage", "middle-aged", "elderly", "child"

Example: `(young male voice)` or `(elderly female voice)`.

### Tone and Pitch Characteristics

Modify the acoustic quality and vocal register using **tone** and **pitch** descriptors:

- **Warmth**: "warm", "bright", "soft", "gentle"
- **Register**: "deep", "high-pitched", "low-pitched", "bright"
- **Timbre**: "nasal", "husky", "smooth", "clear"

Example: `(warm female voice)` or `(deep, husky voice)`.

### Emotional Expression

Shape the affective content through **emotion** keywords:

- "happy", "sad", "excited", "calm"
- "angry", "nervous", "cheerful", "melancholic"
- "smiling", "serious", "playful"

Example: `(cheerful tone)` or `(calm, gentle voice)`.

### Speech Pace and Speed

Control the temporal rhythm using **pace** descriptors:

- "slightly faster", "fast-paced"
- "slow", "very slow", "measured"
- "energetic", "relaxed pace"

Example: `(slightly faster, cheerful tone)` as documented in [`README.md`](https://github.com/OpenBMB/VoxCPM/blob/main/README.md) lines 166-168.

### Stylistic and Delivery Modifiers

Fine-tune the speaking manner through **style** attributes:

- "warm and gentle", "dramatic", "soft"
- "sweet", "smiling", "professional"
- Any free-form phrase conveying vocal quality (e.g., "a slightly husky voice with a playful lilt")

Example: `(young female voice, warm and gentle, slightly smiling)` as shown in [`README.md`](https://github.com/OpenBMB/VoxCPM/blob/main/README.md) line 215.

## Implementation Reference

### CLI Argument Handling

The command-line interface processes the `--control` argument in [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py) (lines 355-364), converting user input into the parenthesized format before inference. The CLI usage follows this pattern:

```bash
voxcpm design \
    --text "Hello world, welcome to VoxCPM!" \
    --control "young female voice, warm and gentle, slightly smiling" \
    --output out.wav

```

### Python API Integration

The core model class in [`src/voxcpm/model/voxcpm.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm.py) exposes the `generate()` method which accepts a `control_instruction` parameter. This parameter feeds directly into the voice-conditioning mechanism without requiring manual string formatting.

## Practical Code Examples

### Python API Usage

Generate speech with controlled attributes using the model's Python interface:

```python
from voxcpm import VoxCPM
import soundfile as sf

model = VoxCPM.from_pretrained("OpenBMB/VoxCPM2", load_denoiser=False)

# Combine multiple attributes: tone, gender, pace, and emotion

text = "Hello, this is a controllable clone."
control = "warm female voice, slightly faster, cheerful tone"

wav = model.generate(text=text, control_instruction=control)
sf.write("controllable.wav", wav, model.tts_model.sample_rate)

```

### Command Line Interface

Rapid prototyping through the CLI allows quick iteration on voice designs:

```bash
voxcpm design \
    --text "Welcome to the future of speech synthesis." \
    --control "deep male voice, professional tone, slightly slower" \
    --output professional_voice.wav

```

### Gradio Web Interface

For interactive voice design, implement a web UI that forwards text prompts to the model:

```python
import gradio as gr
from voxcpm import VoxCPM

model = VoxCPM.from_pretrained("OpenBMB/VoxCPM2")

def synthesize(text, control):
    wav = model.generate(text=text, control_instruction=control)
    return (48000, wav)

iface = gr.Interface(
    fn=synthesize,
    inputs=[
        gr.Textbox(label="Text"),
        gr.Textbox(label="Control Instruction (e.g., 'warm female voice, cheerful')"),
    ],
    outputs=gr.Audio(label="Generated Speech")
)
iface.launch()

```

## Summary

- **Gender and age** control requires phrases like "young male" or "elderly female" according to [`README.md`](https://github.com/OpenBMB/VoxCPM/blob/main/README.md) line 44.
- **Tone and pitch** modifications use descriptors such as "warm", "deep", or "bright".
- **Emotional states** including "happy", "sad", and "excited" steer the affective qualities of synthesis.
- **Pace adjustments** like "slightly faster" or "slow" alter the temporal rhythm without changing content.
- **Stylistic cues** such as "smiling" or "nasal" provide fine-grained control over voice texture.
- The control instruction is processed through `build_final_text()` in [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py) and injected into the model via the `control_instruction` parameter in [`src/voxcpm/model/voxcpm.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm.py).

## Frequently Asked Questions

### What is the exact syntax for controlling voice attributes in VoxCPM?

Wrap your descriptive voice attributes inside parentheses preceding the input text. The CLI utility in [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py) automatically formats this via the `build_final_text()` function, which inserts the control string between parentheses immediately before the content text. For example: `(warm female voice)Hello world`.

### Can I combine multiple voice attributes in a single text prompt?

Yes. VoxCPM supports comma-separated combinations of attributes within the same parenthesized instruction. As demonstrated in the CLI examples from [`README.md`](https://github.com/OpenBMB/VoxCPM/blob/main/README.md) line 215, you can specify `(young female voice, warm and gentle, slightly smiling, slightly faster)` to simultaneously control age, gender, tone, expression, and pace.

### Do I need reference audio clips to use text prompt voice control?

No. VoxCPM supports **zero-shot voice design**, meaning the text prompts alone can generate voices with specified attributes without requiring reference audio files. This distinguishes the control mechanism from voice cloning features that require audio samples.

### Where does the control instruction connect to the model architecture?

The control instruction flows from the CLI/API layer into the model's encoder as a conditioning signal. According to the source implementation, the parenthesized text is treated as a voice-design conditioning token sequence that influences the downstream pipeline components (LocEnc → TSLM → RALM → LocDiT), ultimately steering the acoustic generation without modifying the textual content itself.