How to Deploy VoxCPM for High-Throughput Production with Concurrency
Deploy VoxCPM for high-throughput production using either Nano-vLLM-VoxCPM (recommended) for async batched inference with ~0.13 RTF, or increase Gradio's default_concurrency_limit in app.py for limited parallel processing.
Deploying OpenBMB/VoxCPM in production environments requires choosing between two distinct serving architectures. The standard PyTorch implementation provides simplicity for prototyping, while the specialized Nano-vLLM-VoxCPM engine delivers the concurrency and throughput necessary for production workloads. This guide covers both approaches with specific configuration details extracted from the source code to help you deploy VoxCPM for high-throughput production with concurrency.
Production Deployment Architectures
VoxCPM supports two primary serving modes that differ significantly in concurrency handling and throughput characteristics.
Nano-vLLM-VoxCPM (Recommended)
The Nano-vLLM-VoxCPM engine is a dedicated inference server built on top of the Nano-vLLM framework. According to the source documentation in README.md (lines 245-263), this approach implements asynchronous request batching and GPU pipeline optimization specifically for the VoxCPM diffusion autoregressive pipeline (LocEnc → TSLM → RALM → LocDiT).
Key characteristics include:
- Async batching: Automatically groups incoming requests with a default
max_batch_size=16until a timeout or capacity trigger is reached - Full GPU utilization: Shares a single model instance across concurrent requests using CUDA streams and torch-compile optimizations
- FastAPI backend: Provides container-ready endpoints suitable for Kubernetes or Docker deployments
- Performance: Achieves ~0.13 RTF (real-time factor) on an RTX 4090, scaling linearly with batch size
Standard PyTorch with Gradio
The reference implementation in app.py (lines 889-891) uses Gradio's queue mechanism for request management. By default, this configuration sets default_concurrency_limit=1, meaning requests process sequentially even when multiple workers are available.
Performance characteristics:
- Single-threaded by default: One active inference at a time per GPU
- Memory overhead: Each concurrent worker requires separate model buffer allocations
- RTF: ~0.30 per request on RTX 4090 hardware
- Configuration: Modify
default_concurrency_limitto enable limited parallelism (requires proportional GPU memory increases)
Implementation Guide
High-Throughput Async Server Setup
For production environments requiring maximum throughput, install and configure the Nano-vLLM-VoxCPM engine:
pip install nano-vllm-voxcpm
Deploy the server with automatic batching:
# server_production.py
from nanovllm_voxcpm import VoxCPM
import numpy as np
import soundfile as sf
# Initialize with specific GPU allocation
server = VoxCPM.from_pretrained(
model="openbmb/VoxCPM2", # or local path
devices=[0], # GPU IDs to use
compile=True, # Enable torch.compile for optimization
)
# The server exposes a FastAPI endpoint at POST /generate
# Accepts JSON: {"text": "Hello", "cfg_value": 2.0, "inference_timesteps": 10}
# Returns: {"audio": base64_wav, "sample_rate": 48000}
# Direct generation example for testing
chunks = list(server.generate(target_text="High throughput deployment"))
audio = np.concatenate(chunks)
sf.write("output.wav", audio, 48000)
# Graceful shutdown
server.stop()
The generate method automatically batches multiple concurrent requests, reducing kernel launch overhead across the diffusion steps. The engine implements zero-copy GPU tensor streaming to minimize CPU overhead during audio transmission.
Concurrent Gradio Configuration
When deploying the standard PyTorch implementation, modify the concurrency settings in your Gradio interface. In src/voxcpm/core.py (lines 13-22 and 88-104), the VoxCPM.from_pretrained method loads either VoxCPMModel or VoxCPM2Model instances that can be shared across workers if carefully managed.
# demo_concurrent.py
from voxcpm import VoxCPM
import gradio as gr
# Load model once - shared reference across workers
model = VoxCPM.from_pretrained(
hf_model_id="openbmb/VoxCPM2",
load_denoiser=False, # Disable to save VRAM for concurrent workers
optimize=True,
)
def synthesize(text: str, cfg: float = 2.0, steps: int = 10):
"""Generate speech with the shared model instance."""
wav = model.generate(
text=text,
cfg_value=cfg,
inference_timesteps=steps
)
return wav, model.tts_model.sample_rate
interface = gr.Interface(
fn=synthesize,
inputs=[
gr.Textbox(label="Text"),
gr.Slider(0.1, 10, value=2.0, label="CFG Scale"),
gr.Slider(1, 30, value=10, label="Inference Steps")
],
outputs=gr.Audio(label="Generated Speech"),
)
# Critical: Increase from default 1 to support parallel processing
interface.queue(
max_size=20,
default_concurrency_limit=4 # Adjust based on GPU memory (4-8 for 24GB VRAM)
).launch(server_name="0.0.0.0", server_port=8808)
Memory consideration: Each concurrent worker maintains separate intermediate buffers for the diffusion pipeline. Monitor GPU memory usage when increasing default_concurrency_limit beyond 4 on consumer hardware.
LoRA Voice Customization in Production
Both serving architectures support on-the-fly voice customization via LoRA weights loaded through the load_lora method in core.py. This allows production systems to serve multiple voice personas from a single base model.
# lora_production.py
from voxcpm import VoxCPM
# Load base model with LoRA weights
model = VoxCPM.from_pretrained(
hf_model_id="openbmb/VoxCPM2",
lora_path="checkpoints/custom_voice_lora.pth",
lora_r=8,
lora_alpha=16,
lora_dropout=0.1,
)
# Generate with customized voice
wav = model.generate(
text="(a calm, elderly male voice) Welcome to the production system.",
cfg_value=2.5,
inference_timesteps=12
)
Performance Benchmarks and Scaling
| Architecture | Concurrency Model | RTF (RTX 4090) | Max Throughput |
|---|---|---|---|
| Nano-vLLM-VoxCPM | Async batched (max 16) | ~0.13 | >100 qps |
| PyTorch + Gradio | Worker pool (configurable) | ~0.30 | Limited by VRAM |
The Nano-vLLM implementation achieves approximately 2.3× faster per-request latency while supporting true concurrent processing. The batching mechanism in Nano-vLLM groups diffusion steps across multiple requests, amortizing the computational cost of the LocDiT and RALM modules.
For horizontal scaling, deploy multiple Nano-vLLM-VoxCPM instances behind a load balancer. Each instance can handle 16 concurrent batched requests on a single GPU, with linear scaling across additional GPUs using the devices parameter.
Summary
- Use Nano-vLLM-VoxCPM for production deployments requiring high throughput and concurrency, leveraging its async batching and FastAPI interface.
- Modify
default_concurrency_limitinapp.py(lines 889-891) only when constrained to the standard PyTorch stack, monitoring GPU memory per additional worker. - Initialize
VoxCPM.from_pretrainedonce and share the instance across request handlers to avoid duplicating model weights in memory. - Enable
compile=Truein Nano-vLLM deployments to leverage torch-compile optimizations for the diffusion pipeline. - Load LoRA weights at initialization to support multiple voice personas without separate model instances.
Frequently Asked Questions
What is the maximum number of concurrent requests VoxCPM can handle?
With Nano-vLLM-VoxCPM, the default max_batch_size is 16 concurrent requests per GPU, though this is configurable based on memory constraints. In the standard Gradio implementation, concurrency is limited by the default_concurrency_limit parameter and available VRAM—typically 4-8 concurrent workers on a 24GB GPU before encountering out-of-memory errors during the diffusion steps.
How does LoRA affect production throughput and memory usage?
LoRA (Low-Rank Adaptation) adds minimal overhead to inference latency because the adapter weights are merged into the base model during the forward pass. Memory usage increases slightly with the rank size (configured via lora_r), but the same base model can serve multiple LoRA checkpoints by dynamically loading different adapter weights via the load_lora method without reloading the full pipeline.
Can I deploy VoxCPM on CPU-only servers for production?
The production deployment guides in README.md (lines 245-263) and the core implementation in core.py assume CUDA-capable devices for the diffusion pipeline (LocEnc, TSLM, RALM, LocDiT). CPU inference is not recommended for high-throughput production scenarios due to the computational requirements of the ZipEnhancer denoiser and autoregressive language modeling components.
What is the difference between generate and generate_streaming in production?
The generate method returns complete audio arrays suitable for batched API responses, while generate_streaming yields audio chunks incrementally. For high-throughput production with Nano-vLLM-VoxCPM, generate is preferred as it enables efficient batching across the full diffusion timesteps. Streaming mode is better suited for latency-sensitive interactive applications where immediate playback initiation outweighs throughput optimization.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →