# VoxCPM Ecosystem Tools: A Complete Guide to VoxCPM.cpp, ONNX, and ComfyUI

> Explore the VoxCPM ecosystem with VoxCPM.cpp for fast inference, ONNX for easy deployment, and ComfyUI for node based TTS. Optimize your voice generation workflows today.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: how-to-guide
- Published: 2026-04-10

---

**The VoxCPM ecosystem includes VoxCPM.cpp for high-performance C++ inference, VoxCPM-ONNX for cross-platform deployment, and multiple ComfyUI extensions for node-based TTS workflows—all built atop the core Python model in the OpenBMB/VoxCPM repository.**

The OpenBMB/VoxCPM repository provides more than a Python library; it anchors a growing ecosystem of **VoxCPM ecosystem tools** that enable deployment across diverse environments. These tools share the same core architecture—LocEnc → TSLM → RALM → LocDiT—while offering specialized runtimes for CPU, GPU, edge devices, and visual programming interfaces.

## Core Architecture: The Foundation of All Ecosystem Tools

Every tool in the VoxCPM ecosystem references the canonical implementation located in the main repository. Understanding this foundation is essential before selecting a specific runtime.

### Python Reference Implementation

The `voxcpm` package defines the four-stage pipeline used by all downstream projects. In [`src/voxcpm/model/voxcpm.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm.py), the top-level `VoxCPM` class orchestrates the encoder, language model, and decoder components.

Key source files that external tools import or replicate include:

- [`src/voxcpm/modules/locenc/local_encoder.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/modules/locenc/local_encoder.py) – Implements the local encoder (LocEnc) that converts raw audio into latent tokens
- [`src/voxcpm/modules/locdit/local_dit.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/modules/locdit/local_dit.py) – Contains the diffusion-autoregressive decoder (LocDiT)
- [`src/voxcpm/modules/audiovae/audio_vae_v2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/modules/audiovae/audio_vae_v2.py) – Houses the AudioVAE V2 decoder responsible for 48 kHz super-resolution output
- [`src/voxcpm/cli.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/cli.py) – Provides the command-line interface sub-commands (`design`, `clone`, `batch`) used as reference entry-points

The Python API serves as the reference implementation that the C++, ONNX, and Rust tools either wrap or re-implement.

## VoxCPM.cpp: High-Performance C++ Runtime

**VoxCPM.cpp** provides GGML/GGUF-based inference for CPU, CUDA, or Vulkan backends. This implementation eliminates Python overhead, offering extremely low-memory, high-speed execution suitable for resource-constrained environments.

The tool parses model weights converted to GGUF format and runs the pipeline natively. Conversion scripts provided in the VoxCPM.cpp repository transform the original PyTorch checkpoints into the quantized GGUF format.

```bash

# Convert PyTorch checkpoint to GGUF format

voxcpm-convert --ckpt VoxCPM2.pt --gguf VoxCPM2.gguf

# Run inference on CPU

./voxcpm -m VoxCPM2.gguf -t "A calm, female narrator reads a story." -o story.wav

```

The C++ runtime maintains compatibility with the core model definition while removing the Python dependency entirely, making it ideal for production deployments and edge inference.

## VoxCPM-ONNX: Cross-Platform Export and Inference

**VoxCPM-ONNX** enables export of the model to ONNX format and provides a lightweight ONNX Runtime wrapper for cross-platform CPU inference. This approach bridges the gap between the Python reference and fully native implementations.

The export process uses a script in the main repository (typically located at [`scripts/export_onnx.py`](https://github.com/OpenBMB/VoxCPM/blob/main/scripts/export_onnx.py)) to trace the `VoxCPM` graph with `torch.onnx.export`:

```bash
python -m pip install onnx onnxruntime
python scripts/export_onnx.py \
    --model_dir ./pretrained_models/VoxCPM2 \
    --output voxcpm.onnx

```

Once exported, the ONNX model integrates with the VoxCPM-ONNX runtime wrapper:

```python
import onnxruntime as ort
import numpy as np

sess = ort.InferenceSession("voxcpm.onnx")

# Prepare inputs according to the wrapper's API

audio = sess.run(None, {"text": np.array(["Hello from ONNX!"])})[0]

```

This toolchain is particularly valuable for Windows deployments and environments where Python installation is undesirable but full C++ compilation is impractical.

## ComfyUI Extensions: Visual Workflow Integration

The VoxCPM ecosystem includes multiple node-based integrations for **ComfyUI**, allowing users to embed text-to-speech pipelines inside image-generation workflows without writing code.

### ComfyUI-VoxCPM

The primary ComfyUI extension creates custom nodes that import `voxcpm` directly or load ONNX/GGUF models. Nodes expose parameters such as `text`, `reference_wav`, and `control_prompt`, forwarding data through the standard LocEnc → TSLM → RALM → LocDiT pipeline.

Installation requires placing the package inside the ComfyUI `custom_nodes` folder. Once installed, users can add a **VoxCPM TTS** node, configure the text field, and connect outputs to **Audio Playback** or **Save Audio** nodes.

```json
{
  "type": "VoxCPM_TTS",
  "inputs": {
    "text": "(Cheerful, high-energy voice)Welcome to ComfyUI with VoxCPM!",
    "reference_wav": null
  }
}

```

### ComfyUI-VoxCPMTTS

An alternative extension, **ComfyUI-VoxCPMTTS**, focuses exclusively on TTS-only nodes, providing a streamlined interface for audio generation without image-generation coupling. Both extensions allow seamless chaining with standard ComfyUI image nodes, enabling synchronized audiovisual content creation.

## Additional Community Tools

Beyond the primary trio of VoxCPM.cpp, ONNX, and ComfyUI tools, the ecosystem includes specialized implementations:

- **VoxCPMANE** – Targets the Apple Neural Engine (ANE) for on-device inference on iOS and macOS devices
- **voxcpm_rs** – A Rust re-implementation offering high-performance, safe concurrency for systems programming contexts
- **TTS WebUI** – A browser-based Gradio/Streamlit interface that calls the Python library directly, suitable for rapid prototyping

All these tools import, export, or re-implement the core model definition found in [`src/voxcpm/model/voxcpm.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm.py), ensuring consistent behavior across runtimes.

## Summary

- **VoxCPM.cpp** delivers GGML/GGUF-based C++ inference for CPU, CUDA, and Vulkan, eliminating Python overhead while maintaining the LocEnc → TSLM → RALM → LocDiT pipeline
- **VoxCPM-ONNX** provides cross-platform deployment via ONNX Runtime, using export scripts from the main repository to convert the PyTorch model
- **ComfyUI extensions** (ComfyUI-VoxCPM and ComfyUI-VoxCPMTTS) enable node-based visual programming integration for TTS workflows
- All ecosystem tools reference the canonical Python implementation in [`src/voxcpm/model/voxcpm.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/model/voxcpm.py) and its constituent modules ([`local_encoder.py`](https://github.com/OpenBMB/VoxCPM/blob/main/local_encoder.py), [`local_dit.py`](https://github.com/OpenBMB/VoxCPM/blob/main/local_dit.py), [`audio_vae_v2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/audio_vae_v2.py))
- The ecosystem supports diverse deployment targets including edge devices (VoxCPMANE), systems programming environments (voxcpm_rs), and browser-based UIs

## Frequently Asked Questions

### What is the difference between VoxCPM.cpp and VoxCPM-ONNX?

**VoxCPM.cpp** is a native C++ implementation using the GGML/GGUF format for maximum performance and minimal memory footprint, requiring compilation from source. **VoxCPM-ONNX** exports the model to the standardized ONNX format and runs via ONNX Runtime, offering easier cross-platform deployment without compilation but potentially higher latency than the optimized C++ implementation.

### Can I use the ComfyUI extensions without installing the full Python package?

No, the ComfyUI extensions typically require the `voxcpm` Python package to be installed in the environment, or they must load the exported ONNX/GGUF models. The nodes act as interfaces to the underlying inference engines rather than standalone implementations. Refer to the specific extension repository (ComfyUI-VoxCPM or ComfyUI-VoxCPMTTS) for exact dependency requirements.

### Which ecosystem tool offers the lowest latency for CPU-only inference?

**VoxCPM.cpp** provides the lowest latency for CPU-only scenarios because it uses quantized GGUF weights and the highly optimized GGML inference engine written in C++. The Python reference implementation and ONNX Runtime wrappers introduce additional overhead from the Python interpreter and runtime abstraction layers, respectively.

### Where is the AudioVAE V2 decoder implemented in the source code?

The AudioVAE V2 decoder is implemented in [`src/voxcpm/modules/audiovae/audio_vae_v2.py`](https://github.com/OpenBMB/VoxCPM/blob/main/src/voxcpm/modules/audiovae/audio_vae_v2.py) within the main OpenBMB/VoxCPM repository. This module handles the final super-resolution stage that produces 48 kHz audio output, and it is referenced by all ecosystem tools that maintain full fidelity to the original model architecture.