How to Use the VoxCPM Python API generate() Method for Text-to-Speech Synthesis

The VoxCPM Python API generate() method returns a NumPy float32 waveform array by consuming the first yield from the internal _generate generator with streaming=False, as implemented in src/voxcpm/core.py lines 154-155.

The VoxCPM library from OpenBMB provides a powerful text-to-speech (TTS) synthesis pipeline through its Python API. At the heart of this interface lies the generate() method, which offers a simplified entry point to the diffusion-based audio generation models defined in the VoxCPM and VoxCPM2 classes. This guide explains how to use the method effectively, from basic invocation to advanced voice cloning scenarios.

Method Signature and Return Value

According to the source code in src/voxcpm/core.py, the generate() method accepts a comprehensive set of arguments through *args and **kwargs and returns a single NumPy ndarray representing the synthesized audio waveform.

Specifically, lines 154-155 define the method as a thin wrapper:

def generate(self, *args, **kwargs) -> np.ndarray:
    return next(self._generate(*args, streaming=False, **kwargs))

This implementation delegates all processing to the private _generate method while forcing streaming=False, ensuring the function returns the complete waveform rather than a generator object.

The Internal Synthesis Pipeline

Understanding the generate() method requires examining the underlying _generate workflow in src/voxcpm/core.py. The pipeline executes several distinct phases:

Parameter Validation and Setup

The _generate method beginning at line 160 validates all incoming arguments, including text content, optional prompt audio paths, reference audio for voice cloning, CFG scale, inference steps, and length limits. Lines 160-176 enforce constraints such as non-empty text strings and validate that provided file paths exist on disk. Mutual dependency checks ensure that prompt_wav_path and prompt_text are supplied together when using continuation mode.

Optional Preprocessing

When enabled, the pipeline applies two preprocessing stages defined in the early sections of core.py:

Audio Denoising (lines 29-38): If denoise=True and a denoiser model is loaded (implemented in src/voxcpm/zipenhancer.py), the method processes prompt and reference audio files through the denoiser, creating temporary cleaned versions before feeding them to the model.

Text Normalization (lines 56-62): When normalize=True, the method lazily instantiates a TextNormalizer (defined in src/voxcpm/utils/text_normalize.py) to preprocess input text, handling punctuation and formatting standardization.

Prompt Cache Construction

For continuation or voice-cloning scenarios, lines 41-53 build a prompt cache using the underlying TTS model's build_prompt_cache() API. VoxCPM2 models additionally accept a reference wav path at this stage, enabling zero-shot voice cloning capabilities implemented in src/voxcpm/model/voxcpm2.py.

Model Inference and Output

Lines 63-73 hand the prepared arguments to the model's private _generate_with_prompt_cache method (located in src/voxcpm/model/voxcpm.py), which executes the diffusion inference loop. Finally, lines 76-78 yield the resulting waveform, which the public generate() method consumes and returns as a one-dimensional NumPy float32 array.

Basic Usage Example

To synthesize speech, first instantiate the VoxCPM class using from_pretrained(), then call generate() with your target text:

from voxcpm import VoxCPM

# Load the pretrained model (will download if not cached)

vc = VoxCPM.from_pretrained(
    hf_model_id="OpenBMB/VoxCPM",
    load_denoiser=False,   # optional denoiser

)

# Simple text-to-speech

waveform = vc.generate(text="Welcome to Vox CPM!")

# `waveform` is a NumPy array; you can save it with soundfile or play it directly.

Advanced Features and Usage Modes

The generate() method supports several specialized operating modes through specific parameter combinations:

Audio Continuation

To continue existing audio with matching text, provide both the audio file and the text prefix that corresponds to it:

waveform = vc.generate(
    text="and we will continue the story.",
    prompt_wav_path="samples/intro.wav",
    prompt_text="Once upon a time,",
)

This triggers the prompt cache construction path at lines 41-53 of core.py, conditioning the model on the acoustic features of the provided sample.

Voice Cloning with VoxCPM-2

For voice cloning capabilities, supply a reference audio file. This feature requires the VoxCPM2 model architecture implemented in src/voxcpm/model/voxcpm2.py:

waveform = vc.generate(
    text="I sound like you now.",
    reference_wav_path="samples/voice_demo.wav",
)

Streaming Generation

While generate() returns complete waveforms, the library provides generate_streaming() for low-latency applications. This method yields intermediate audio chunks suitable for real-time playback:

for chunk in vc.generate_streaming(text="Streaming synthesis example"):
    # `chunk` is a NumPy array for the current generation step

    play(chunk)   # replace with your audio playback routine

Denoising and Text Normalization

Combine audio preprocessing and text cleaning in a single call:

waveform = vc.generate(
    text="  This   text   needs   cleaning! ",
    prompt_wav_path="noisy_prompt.wav",
    denoise=True,
    normalize=True,
)

When denoise=True, the system utilizes the ZipEnhancer model from src/voxcpm/zipenhancer.py to clean input audio before processing.

Summary

  • Simple wrapper: The generate() method returns the first yield from _generate(*args, streaming=False, **kwargs) as a NumPy array (lines 154-155 in src/voxcpm/core.py).
  • Rich parameters: Supports text prompts, audio continuation via prompt_wav_path, voice cloning via reference_wav_path, CFG scaling, and inference step control.
  • Preprocessing options: Optional denoising via src/voxcpm/zipenhancer.py and text normalization via src/voxcpm/utils/text_normalize.py.
  • Model backends: Delegates to VoxCPMModel or VoxCPM2Model classes in the src/voxcpm/model/ directory for actual diffusion inference.

Frequently Asked Questions

What is the difference between generate() and generate_streaming()?

The generate() method calls the internal _generate function with streaming=False and returns the first yielded result as a complete NumPy array, suitable for batch processing or file export. In contrast, generate_streaming() sets streaming=True, yielding partial audio chunks during the diffusion inference process for real-time playback applications.

How do I enable voice cloning with VoxCPM?

Voice cloning requires using the VoxCPM-2 model architecture and providing a reference_wav_path parameter to the generate() method. The reference audio is processed through the prompt cache construction logic (lines 41-53 in src/voxcpm/core.py) and the VoxCPM2Model class in src/voxcpm/model/voxcpm2.py to extract speaker characteristics before generation.

What audio format does the generate() method return?

The method returns a one-dimensional NumPy float32 array containing the raw audio waveform samples. This standard format allows direct integration with audio processing libraries such as soundfile, scipy.io.wavfile, or real-time playback systems. The sample rate depends on the specific pretrained model configuration loaded via from_pretrained().

Why am I getting validation errors when using prompt_wav_path?

The validation logic at lines 160-176 in src/voxcpm/core.py requires that prompt_wav_path be accompanied by prompt_text when using continuation mode. Additionally, the method checks that the specified file path exists and that the text parameter is non-empty. Ensure both parameters are provided as valid strings and that the audio file exists on disk.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →