CLI Commands for Batch Processing TTS Requests in VoxCPM: A Complete Guide
The voxcpm batch command reads a text file containing one utterance per line and generates separate WAV files for each entry, supporting voice cloning, continuation prompts, and voice design controls.
VoxCPM is an open-source text-to-speech system that provides robust command-line tools for synthesizing speech at scale. The CLI includes a dedicated batch sub-command designed specifically for processing multiple TTS requests efficiently without requiring repeated model loading. This guide covers the complete implementation details, argument specifications, and practical examples for batch processing in the OpenBMB/VoxCPM repository.
How the Batch Command Works
The batch processing pipeline in src/voxcpm/cli.py follows a structured execution flow that validates inputs once and reuses loaded models across all utterances.
Argument parsing begins in _build_parser() at lines 94-105, where the batch sub-parser is configured with required and optional flags. The cmd_batch function first validates that the input file exists and contains at least one non-empty line (lines 89-99).
Model loading utilizes the same load_model logic used for single-sample commands (lines 176-236), ensuring consistent initialization across different CLI modes. This prevents the overhead of reloading weights between utterances.
Optional audio inputs are validated at lines 103-114. If --prompt-audio or --reference-audio paths are provided, the system verifies their existence before passing them to the synthesis pipeline.
The generation loop (lines 117-138) processes each line individually by calling build_final_text, generating the audio array through the model, and saving results as output_<index>.wav in the specified output directory. Finally, reporting at lines 140-143 prints a summary of successful versus total generations to the console.
Required Arguments and Options
The voxcpm batch command requires specific paths while offering extensive customization through optional flags documented in the parser's epilog (lines 660-667).
Required arguments:
--inputor-i: Path to a plain-text file where each line represents a separate TTS request--output-diror-od: Directory where generated WAV files will be written
Voice control options:
--control: Voice-design instruction passed to VoxCPM2 (e.g., "warm female voice, friendly tone")--reference-audio: Path to reference speaker audio for voice cloning (VoxCPM2 only)--prompt-audio,--prompt-text,--prompt-file: Enable continuation mode with reference audio and transcript
Generation parameters:
--cfg-value: CFG guidance scale (default 2.0)--inference-timesteps: Number of inference steps (default 10)--normalize: Enable text normalization preprocessing--denoise: Apply speech enhancement to prompt/reference audio
Model configuration:
--model-path,--hf-model-id,--cache-dir: Standard model selection and caching options shared with single-sample commands
Practical Batch Processing Examples
Simple Batch Generation
Create a text file with one utterance per line:
# texts.txt
Hello world
How are you today?
Welcome to VoxCPM batch demo
Execute batch synthesis:
voxcpm batch \
--input texts.txt \
--output-dir ./batch_outputs \
--cfg-value 2.5 \
--inference-timesteps 12
This creates ./batch_outputs/output_001.wav, output_002.wav, and output_003.wav, each containing synthesized speech for the corresponding line.
Batch with Voice Design
Apply consistent stylistic controls across all outputs:
voxcpm batch \
--input texts.txt \
--output-dir ./warm_female \
--control "warm female voice, friendly tone" \
--cfg-value 3.0
Batch Voice Cloning (VoxCPM2 Only)
Clone a specific speaker across multiple utterances:
voxcpm batch \
--input texts.txt \
--output-dir ./cloned_speaker \
--reference-audio speaker_ref.wav \
--cfg-value 2.0
Batch with Continuation Prompts
Generate speech that continues naturally from a reference audio:
voxcpm batch \
--input texts.txt \
--output-dir ./continuation_demo \
--prompt-audio prompt.wav \
--prompt-text "This is the opening line of a story." \
--cfg-value 2.2
Key Implementation Files
The batch functionality spans several critical components in the OpenBMB/VoxCPM repository:
src/voxcpm/cli.py: Implements the full command-line interface including thebatchsub-command, argument validation, and the generation loop (lines 89-143)src/voxcpm/core.py: Contains the coreVoxCPMclass that wraps the TTS, design, and cloning pipelinessrc/voxcpm/model/voxcpm.py: Defines the model architecture and thegeneratemethod invoked during batch processingtests/test_cli.py: Provides unit tests covering CLI parsing and batch command behavior
Summary
- The
voxcpm batchcommand efficiently processes multiple TTS requests from a single text file, generatingoutput_<index>.wavfiles in the specified directory - Input validation occurs once at startup in
cmd_batch(lines 89-99), checking file existence and content before model loading - Models load once via
load_model(lines 176-236) and are reused across all utterances to optimize performance - The command supports advanced features including voice cloning (
--reference-audio), continuation prompts (--prompt-audio), and voice design controls (--control) - Processing reports indicate successful versus total generation counts upon completion (lines 140-143)
- All arguments are documented in the parser epilog at lines 660-667 of
cli.py
Frequently Asked Questions
What file format does the input text file need for batch processing?
The input file must be a plain text file with one utterance per line. The cmd_batch function validates that the file exists and contains at least one non-empty line before processing begins. Empty lines are handled gracefully during the generation loop at lines 117-138.
Can I use different models or voices for each line in a batch file?
No, the current implementation in cli.py loads the model once (lines 176-236) and applies the same configuration—including voice cloning references and design controls—to every line in the batch. To use different voices or models, separate the utterances into different batch files and run the command multiple times.
How does batch processing handle errors for individual utterances?
According to the generation loop implementation (lines 117-138), the batch process continues independently for each line. The final reporting mechanism at lines 140-143 tracks successful generations versus total attempts, allowing you to identify which specific outputs succeeded without stopping the entire batch on a single failure.
Where can I find the complete list of available options for the batch command?
The complete argument specification is documented in the parser's help text, accessible via voxcpm batch --help or voxcpm -h. The epilog containing this documentation is defined at lines 660-667 in src/voxcpm/cli.py, while the argument parsers are constructed in _build_parser() at lines 94-105.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →