# How to Deploy VoxCPM for High-Throughput Production with Concurrency

> Learn to deploy VoxCPM for high-throughput production with concurrency. Explore Nano-vLLM-VoxCPM for async batched inference or adjust Gradio settings for parallel processing.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: how-to-guide
- Published: 2026-04-10

---

**Deploy VoxCPM for high-throughput production using either Nano-vLLM-VoxCPM (recommended) for async batched inference with ~0.13 RTF, or increase Gradio's `default_concurrency_limit` in [`app.py`](https://github.com/OpenBMB/VoxCPM/blob/main/app.py) for limited parallel processing.**

Deploying OpenBMB/VoxCPM in production environments requires choosing between two distinct serving architectures. The standard PyTorch implementation provides simplicity for prototyping, while the specialized Nano-vLLM-VoxCPM engine delivers the concurrency and throughput necessary for production workloads. This guide covers both approaches with specific configuration details extracted from the source code to help you deploy VoxCPM for high-throughput production with concurrency.

## Production Deployment Architectures

VoxCPM supports two primary serving modes that differ significantly in concurrency handling and throughput characteristics.

### Nano-vLLM-VoxCPM (Recommended)

The **Nano-vLLM-VoxCPM** engine is a dedicated inference server built on top of the Nano-vLLM framework. According to the source documentation in [`README.md`](https://github.com/OpenBMB/VoxCPM/blob/main/README.md) (lines 245-263), this approach implements asynchronous request batching and GPU pipeline optimization specifically for the VoxCPM diffusion autoregressive pipeline (LocEnc → TSLM → RALM → LocDiT).

Key characteristics include:
- **Async batching**: Automatically groups incoming requests with a default `max_batch_size=16` until a timeout or capacity trigger is reached
- **Full GPU utilization**: Shares a single model instance across concurrent requests using CUDA streams and torch-compile optimizations
- **FastAPI backend**: Provides container-ready endpoints suitable for Kubernetes or Docker deployments
- **Performance**: Achieves ~0.13 RTF (real-time factor) on an RTX 4090, scaling linearly with batch size

### Standard PyTorch with Gradio

The reference implementation in [`app.py`](https://github.com/OpenBMB/VoxCPM/blob/main/app.py) (lines 889-891) uses Gradio's queue mechanism for request management. By default, this configuration sets `default_concurrency_limit=1`, meaning requests process sequentially even when multiple workers are available.

Performance characteristics:
- **Single-threaded by default**: One active inference at a time per GPU
- **Memory overhead**: Each concurrent worker requires separate model buffer allocations
- **RTF**: ~0.30 per request on RTX 4090 hardware
- **Configuration**: Modify `default_concurrency_limit` to enable limited parallelism (requires proportional GPU memory increases)

## Implementation Guide

### High-Throughput Async Server Setup

For production environments requiring maximum throughput, install and configure the Nano-vLLM-VoxCPM engine:

```bash
pip install nano-vllm-voxcpm

```

Deploy the server with automatic batching:

```python

# server_production.py

from nanovllm_voxcpm import VoxCPM
import numpy as np
import soundfile as sf

# Initialize with specific GPU allocation

server = VoxCPM.from_pretrained(
    model="openbmb/VoxCPM2",  # or local path

    devices=[0],               # GPU IDs to use

    compile=True,              # Enable torch.compile for optimization

)

# The server exposes a FastAPI endpoint at POST /generate

# Accepts JSON: {"text": "Hello", "cfg_value": 2.0, "inference_timesteps": 10}

# Returns: {"audio": base64_wav, "sample_rate": 48000}

# Direct generation example for testing

chunks = list(server.generate(target_text="High throughput deployment"))
audio = np.concatenate(chunks)
sf.write("output.wav", audio, 48000)

# Graceful shutdown

server.stop()

```

The `generate` method automatically batches multiple concurrent requests, reducing kernel launch overhead across the diffusion steps. The engine implements zero-copy GPU tensor streaming to minimize CPU overhead during audio transmission.

### Concurrent Gradio Configuration

When deploying the standard PyTorch implementation, modify the concurrency settings in your Gradio interface. In [`src/voxcpm/core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/core.py) (lines 13-22 and 88-104), the `VoxCPM.from_pretrained` method loads either `VoxCPMModel` or `VoxCPM2Model` instances that can be shared across workers if carefully managed.

```python

# demo_concurrent.py

from voxcpm import VoxCPM
import gradio as gr

# Load model once - shared reference across workers

model = VoxCPM.from_pretrained(
    hf_model_id="openbmb/VoxCPM2",
    load_denoiser=False,    # Disable to save VRAM for concurrent workers

    optimize=True,
)

def synthesize(text: str, cfg: float = 2.0, steps: int = 10):
    """Generate speech with the shared model instance."""
    wav = model.generate(
        text=text,
        cfg_value=cfg,
        inference_timesteps=steps
    )
    return wav, model.tts_model.sample_rate

interface = gr.Interface(
    fn=synthesize,
    inputs=[
        gr.Textbox(label="Text"),
        gr.Slider(0.1, 10, value=2.0, label="CFG Scale"),
        gr.Slider(1, 30, value=10, label="Inference Steps")
    ],
    outputs=gr.Audio(label="Generated Speech"),
)

# Critical: Increase from default 1 to support parallel processing

interface.queue(
    max_size=20,
    default_concurrency_limit=4  # Adjust based on GPU memory (4-8 for 24GB VRAM)

).launch(server_name="0.0.0.0", server_port=8808)

```

**Memory consideration**: Each concurrent worker maintains separate intermediate buffers for the diffusion pipeline. Monitor GPU memory usage when increasing `default_concurrency_limit` beyond 4 on consumer hardware.

### LoRA Voice Customization in Production

Both serving architectures support on-the-fly voice customization via LoRA weights loaded through the `load_lora` method in [`core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/core.py). This allows production systems to serve multiple voice personas from a single base model.

```python

# lora_production.py

from voxcpm import VoxCPM

# Load base model with LoRA weights

model = VoxCPM.from_pretrained(
    hf_model_id="openbmb/VoxCPM2",
    lora_path="checkpoints/custom_voice_lora.pth",
    lora_r=8,
    lora_alpha=16,
    lora_dropout=0.1,
)

# Generate with customized voice

wav = model.generate(
    text="(a calm, elderly male voice) Welcome to the production system.",
    cfg_value=2.5,
    inference_timesteps=12
)

```

## Performance Benchmarks and Scaling

| Architecture | Concurrency Model | RTF (RTX 4090) | Max Throughput |
|-------------|-------------------|----------------|----------------|
| **Nano-vLLM-VoxCPM** | Async batched (max 16) | ~0.13 | >100 qps |
| **PyTorch + Gradio** | Worker pool (configurable) | ~0.30 | Limited by VRAM |

The Nano-vLLM implementation achieves approximately **2.3× faster** per-request latency while supporting true concurrent processing. The batching mechanism in Nano-vLLM groups diffusion steps across multiple requests, amortizing the computational cost of the LocDiT and RALM modules.

For horizontal scaling, deploy multiple Nano-vLLM-VoxCPM instances behind a load balancer. Each instance can handle 16 concurrent batched requests on a single GPU, with linear scaling across additional GPUs using the `devices` parameter.

## Summary

- **Use Nano-vLLM-VoxCPM** for production deployments requiring high throughput and concurrency, leveraging its async batching and FastAPI interface.
- **Modify `default_concurrency_limit`** in [`app.py`](https://github.com/OpenBMB/VoxCPM/blob/main/app.py) (lines 889-891) only when constrained to the standard PyTorch stack, monitoring GPU memory per additional worker.
- **Initialize `VoxCPM.from_pretrained`** once and share the instance across request handlers to avoid duplicating model weights in memory.
- **Enable `compile=True`** in Nano-vLLM deployments to leverage torch-compile optimizations for the diffusion pipeline.
- **Load LoRA weights at initialization** to support multiple voice personas without separate model instances.

## Frequently Asked Questions

### What is the maximum number of concurrent requests VoxCPM can handle?

With Nano-vLLM-VoxCPM, the default `max_batch_size` is 16 concurrent requests per GPU, though this is configurable based on memory constraints. In the standard Gradio implementation, concurrency is limited by the `default_concurrency_limit` parameter and available VRAM—typically 4-8 concurrent workers on a 24GB GPU before encountering out-of-memory errors during the diffusion steps.

### How does LoRA affect production throughput and memory usage?

LoRA (Low-Rank Adaptation) adds minimal overhead to inference latency because the adapter weights are merged into the base model during the forward pass. Memory usage increases slightly with the rank size (configured via `lora_r`), but the same base model can serve multiple LoRA checkpoints by dynamically loading different adapter weights via the `load_lora` method without reloading the full pipeline.

### Can I deploy VoxCPM on CPU-only servers for production?

The production deployment guides in [`README.md`](https://github.com/OpenBMB/VoxCPM/blob/main/README.md) (lines 245-263) and the core implementation in [`core.py`](https://github.com/OpenBMB/VoxCPM/blob/main/core.py) assume CUDA-capable devices for the diffusion pipeline (LocEnc, TSLM, RALM, LocDiT). CPU inference is not recommended for high-throughput production scenarios due to the computational requirements of the ZipEnhancer denoiser and autoregressive language modeling components.

### What is the difference between `generate` and `generate_streaming` in production?

The `generate` method returns complete audio arrays suitable for batched API responses, while `generate_streaming` yields audio chunks incrementally. For high-throughput production with Nano-vLLM-VoxCPM, `generate` is preferred as it enables efficient batching across the full diffusion timesteps. Streaming mode is better suited for latency-sensitive interactive applications where immediate playback initiation outweighs throughput optimization.