# How to Create Custom Voices from Text Descriptions in VoxCPM Voice Design

> Learn to create custom voices from text descriptions with VoxCPM Voice Design. Generate novel voices using control instructions and the VoxCPM2 model without reference audio.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: how-to-guide
- Published: 2026-04-10

---

**VoxCPM Voice Design generates custom voices from text descriptions by prepending control instructions inside parentheses to the target text, allowing the VoxCPM2 model to synthesize novel voices without reference audio.**

VoxCPM is an open-source text-to-speech framework by OpenBMB that supports voice cloning and zero-shot voice creation. Unlike traditional voice cloning that requires a reference audio sample, VoxCPM Voice Design creates entirely new voices from natural language descriptions such as "warm female voice" or "young surfer dude."

## How Voice Design Works

The Voice Design feature relies on a **Control Instruction** mechanism. When you provide a text description of the desired voice characteristics, the system wraps this description in parentheses and concatenates it with your target text. This combined string acts as a conditioning signal for the latent diffusion model.

The key innovation is that VoxCPM2 was trained on paired control-instruction and audio examples. During inference, the model interprets the text inside the parentheses as acoustic conditioning, extrapolating from learned patterns to generate a voice that matches the description—even if that specific voice never appeared in the training data.

## Implementation Architecture

### CLI Parsing

In [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py), the `--control` argument is processed by the `build_final_text` function at lines 71-73:

```python
def build_final_text(text: str, control: str | None) -> str:
    control = (control or "").strip()
    return f"({control}){text}" if control else text

```

This function transforms a control instruction like "warm female voice" and target text "Hello world" into the formatted string `"(warm female voice)Hello world"`.

### Core Generation Pipeline

The formatted text flows into [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py), where the `VoxCPM.generate` method handles the synthesis. At lines 55-63, the code calls:

```python
generate_result = self.tts_model._generate_with_prompt_cache(
    target_text=text,
    prompt_cache=fixed_prompt_cache,
    # ...

)

```

When `fixed_prompt_cache` is `None` (indicating no reference audio), the model treats the entire formatted string—including the control token—as the synthesis prompt.

### Model Conditioning

Inside [`src/voxcpm/model/voxcpm2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm2.py), the model parses the input sequence and detects the `(<any text>)` pattern. This control token routes through a dedicated **control encoder** that influences both the latent diffusion (LocDiT) and language model streams. The conditioning shapes every generation step, determining the final timbre, pitch, and speaking style of the output audio.

## Practical Usage Examples

### Command-Line Interface

The fastest way to create a custom voice is using the CLI:

```bash
voxcpm design \
    --text "Hello world, this is a custom voice." \
    --control "warm female voice" \
    --output warm_female.wav

```

This command invokes `build_final_text` internally, formats the control instruction, and executes the generation pipeline. The output `warm_female.wav` contains a newly synthesized voice matching your description.

### Python API Integration

For programmatic control, use the Python API directly:

```python
from voxcpm.core import VoxCPM

# Initialize the model

model = VoxCPM.from_pretrained(
    hf_model_id="openbmb/VoxCPM2",
    load_denoiser=False,
)

# Format the control instruction

def build_final_text(text: str, control: str | None) -> str:
    control = (control or "").strip()
    return f"({control}){text}" if control else text

control = "young surfer dude, relaxed, slightly nasal"
target = "Hey, catch that wave! It's a perfect day."
final_text = build_final_text(target, control)

# Generate audio

audio = model.generate(
    text=final_text,
    cfg_value=2.5,
    inference_timesteps=12,
)

# Save output

import soundfile as sf
sf.write("surfer.wav", audio, samplerate=model.tts_model.sample_rate)

```

This approach gives you fine-grained control over diffusion parameters like `cfg_value` (guidance strength) and `inference_timesteps`.

### Gradio Web Interface

To use the graphical interface:

1. Launch the server with `python app.py`
2. Navigate to `http://localhost:7860`
3. Enter your **Control Instruction** (e.g., "A soft, sweet teenage girl")
4. Input your **Target Text**
5. Click **Generate Speech**

The Gradio UI in [`app.py`](https://github.com/OpenBMB/VoxCPM/blob/main/app.py) maps the control field to the same `build_final_text` logic used by the CLI, ensuring consistent behavior across interfaces.

## Key Source Files

- **[`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py)** – Parses the `--control` argument and formats instructions via `build_final_text` (lines 71-73)
- **[`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py)** – Contains the `VoxCPM` class and generation pipeline entry point (lines 55-63)
- **[`src/voxcpm/model/voxcpm2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm2.py)** – Implements the model architecture that processes control tokens during diffusion
- **[`app.py`](https://github.com/OpenBMB/VoxCPM/blob/main/app.py)** – Gradio interface exposing the Voice Design workflow
- **[`tests/test_cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/tests/test_cli.py)** – Unit tests verifying control argument embedding (lines 154-160)

## Summary

- VoxCPM Voice Design creates custom voices from text descriptions without reference audio
- Control instructions use the format `(<description>)<target_text>` via the `build_final_text` function in [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py)
- The VoxCPM2 model interprets parenthetical control tokens as conditioning signals for its latent diffusion architecture
- You can access this functionality through the CLI, Python API, or Gradio web interface

## Frequently Asked Questions

### What is the exact syntax for control instructions in VoxCPM?

Wrap your voice description in parentheses immediately preceding the text you want spoken. For example: `"(deep male voice)Welcome to the system."` The `build_final_text` function in [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py) handles this formatting automatically when you use the `--control` argument.

### Can I combine voice design with voice cloning?

No, these are mutually exclusive modes. When you provide a control instruction, `fixed_prompt_cache` is set to `None` in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py), forcing the model to generate a voice from text description only. To clone a specific speaker, you must omit the control instruction and provide reference audio instead.

### How detailed should the text description be?

VoxCPM2 responds to natural language descriptions including gender, age, emotional tone, and speaking style. Descriptions like "warm female voice" or "young surfer dude, relaxed, slightly nasal" work effectively. The model extrapolates acoustic characteristics from the training data associations with these descriptive phrases.

### Where is the control token processing implemented in the source code?

The control token parsing occurs in [`src/voxcpm/model/voxcpm2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm2.py), where the model detects the parenthetical pattern and routes it through the control encoder. This encoder influences both the LocDiT diffusion process and the language model components to shape the final audio output.