# How to Use the VoxCPM Python API generate() Method for Text-to-Speech Synthesis

> Learn how to use the VoxCPM Python API generate() method for text to speech. Get a NumPy float32 waveform array for seamless audio synthesis. Explore the core functionality.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: how-to-guide
- Published: 2026-04-10

---

**The VoxCPM Python API `generate()` method returns a NumPy `float32` waveform array by consuming the first yield from the internal `_generate` generator with `streaming=False`, as implemented in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py) lines 154-155.**

The **VoxCPM** library from OpenBMB provides a powerful text-to-speech (TTS) synthesis pipeline through its Python API. At the heart of this interface lies the `generate()` method, which offers a simplified entry point to the diffusion-based audio generation models defined in the `VoxCPM` and `VoxCPM2` classes. This guide explains how to use the method effectively, from basic invocation to advanced voice cloning scenarios.

## Method Signature and Return Value

According to the source code in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py), the `generate()` method accepts a comprehensive set of arguments through `*args` and `**kwargs` and returns a single NumPy ndarray representing the synthesized audio waveform.

Specifically, lines 154-155 define the method as a thin wrapper:

```python
def generate(self, *args, **kwargs) -> np.ndarray:
    return next(self._generate(*args, streaming=False, **kwargs))

```

This implementation delegates all processing to the private `_generate` method while forcing `streaming=False`, ensuring the function returns the complete waveform rather than a generator object.

## The Internal Synthesis Pipeline

Understanding the `generate()` method requires examining the underlying `_generate` workflow in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py). The pipeline executes several distinct phases:

### Parameter Validation and Setup

The `_generate` method beginning at line 160 validates all incoming arguments, including text content, optional prompt audio paths, reference audio for voice cloning, CFG scale, inference steps, and length limits. Lines 160-176 enforce constraints such as non-empty text strings and validate that provided file paths exist on disk. Mutual dependency checks ensure that `prompt_wav_path` and `prompt_text` are supplied together when using continuation mode.

### Optional Preprocessing

When enabled, the pipeline applies two preprocessing stages defined in the early sections of [`core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/core.py):

**Audio Denoising** (lines 29-38): If `denoise=True` and a denoiser model is loaded (implemented in [`src/voxcpm/zipenhancer.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/zipenhancer.py)), the method processes prompt and reference audio files through the denoiser, creating temporary cleaned versions before feeding them to the model.

**Text Normalization** (lines 56-62): When `normalize=True`, the method lazily instantiates a `TextNormalizer` (defined in [`src/voxcpm/utils/text_normalize.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/utils/text_normalize.py)) to preprocess input text, handling punctuation and formatting standardization.

### Prompt Cache Construction

For continuation or voice-cloning scenarios, lines 41-53 build a prompt cache using the underlying TTS model's `build_prompt_cache()` API. VoxCPM2 models additionally accept a reference wav path at this stage, enabling zero-shot voice cloning capabilities implemented in [`src/voxcpm/model/voxcpm2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm2.py).

### Model Inference and Output

Lines 63-73 hand the prepared arguments to the model's private `_generate_with_prompt_cache` method (located in [`src/voxcpm/model/voxcpm.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm.py)), which executes the diffusion inference loop. Finally, lines 76-78 yield the resulting waveform, which the public `generate()` method consumes and returns as a one-dimensional NumPy `float32` array.

## Basic Usage Example

To synthesize speech, first instantiate the `VoxCPM` class using `from_pretrained()`, then call `generate()` with your target text:

```python
from voxcpm import VoxCPM

# Load the pretrained model (will download if not cached)

vc = VoxCPM.from_pretrained(
    hf_model_id="OpenBMB/VoxCPM",
    load_denoiser=False,   # optional denoiser

)

# Simple text-to-speech

waveform = vc.generate(text="Welcome to Vox CPM!")

# `waveform` is a NumPy array; you can save it with soundfile or play it directly.

```

## Advanced Features and Usage Modes

The `generate()` method supports several specialized operating modes through specific parameter combinations:

### Audio Continuation

To continue existing audio with matching text, provide both the audio file and the text prefix that corresponds to it:

```python
waveform = vc.generate(
    text="and we will continue the story.",
    prompt_wav_path="samples/intro.wav",
    prompt_text="Once upon a time,",
)

```

This triggers the prompt cache construction path at lines 41-53 of [`core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/core.py), conditioning the model on the acoustic features of the provided sample.

### Voice Cloning with VoxCPM-2

For voice cloning capabilities, supply a reference audio file. This feature requires the VoxCPM2 model architecture implemented in [`src/voxcpm/model/voxcpm2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm2.py):

```python
waveform = vc.generate(
    text="I sound like you now.",
    reference_wav_path="samples/voice_demo.wav",
)

```

### Streaming Generation

While `generate()` returns complete waveforms, the library provides `generate_streaming()` for low-latency applications. This method yields intermediate audio chunks suitable for real-time playback:

```python
for chunk in vc.generate_streaming(text="Streaming synthesis example"):
    # `chunk` is a NumPy array for the current generation step

    play(chunk)   # replace with your audio playback routine

```

### Denoising and Text Normalization

Combine audio preprocessing and text cleaning in a single call:

```python
waveform = vc.generate(
    text="  This   text   needs   cleaning! ",
    prompt_wav_path="noisy_prompt.wav",
    denoise=True,
    normalize=True,
)

```

When `denoise=True`, the system utilizes the ZipEnhancer model from [`src/voxcpm/zipenhancer.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/zipenhancer.py) to clean input audio before processing.

## Summary

- **Simple wrapper**: The `generate()` method returns the first yield from `_generate(*args, streaming=False, **kwargs)` as a NumPy array (lines 154-155 in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py)).
- **Rich parameters**: Supports text prompts, audio continuation via `prompt_wav_path`, voice cloning via `reference_wav_path`, CFG scaling, and inference step control.
- **Preprocessing options**: Optional denoising via [`src/voxcpm/zipenhancer.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/zipenhancer.py) and text normalization via [`src/voxcpm/utils/text_normalize.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/utils/text_normalize.py).
- **Model backends**: Delegates to `VoxCPMModel` or `VoxCPM2Model` classes in the `src/voxcpm/model/` directory for actual diffusion inference.

## Frequently Asked Questions

### What is the difference between `generate()` and `generate_streaming()`?

The `generate()` method calls the internal `_generate` function with `streaming=False` and returns the first yielded result as a complete NumPy array, suitable for batch processing or file export. In contrast, `generate_streaming()` sets `streaming=True`, yielding partial audio chunks during the diffusion inference process for real-time playback applications.

### How do I enable voice cloning with VoxCPM?

Voice cloning requires using the **VoxCPM-2** model architecture and providing a `reference_wav_path` parameter to the `generate()` method. The reference audio is processed through the prompt cache construction logic (lines 41-53 in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py)) and the `VoxCPM2Model` class in [`src/voxcpm/model/voxcpm2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm2.py) to extract speaker characteristics before generation.

### What audio format does the `generate()` method return?

The method returns a one-dimensional NumPy `float32` array containing the raw audio waveform samples. This standard format allows direct integration with audio processing libraries such as `soundfile`, `scipy.io.wavfile`, or real-time playback systems. The sample rate depends on the specific pretrained model configuration loaded via `from_pretrained()`.

### Why am I getting validation errors when using `prompt_wav_path`?

The validation logic at lines 160-176 in [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py) requires that `prompt_wav_path` be accompanied by `prompt_text` when using continuation mode. Additionally, the method checks that the specified file path exists and that the text parameter is non-empty. Ensure both parameters are provided as valid strings and that the audio file exists on disk.