VoxCPM Ecosystem Tools: A Complete Guide to VoxCPM.cpp, ONNX, and ComfyUI

The VoxCPM ecosystem includes VoxCPM.cpp for high-performance C++ inference, VoxCPM-ONNX for cross-platform deployment, and multiple ComfyUI extensions for node-based TTS workflows—all built atop the core Python model in the OpenBMB/VoxCPM repository.

The OpenBMB/VoxCPM repository provides more than a Python library; it anchors a growing ecosystem of VoxCPM ecosystem tools that enable deployment across diverse environments. These tools share the same core architecture—LocEnc → TSLM → RALM → LocDiT—while offering specialized runtimes for CPU, GPU, edge devices, and visual programming interfaces.

Core Architecture: The Foundation of All Ecosystem Tools

Every tool in the VoxCPM ecosystem references the canonical implementation located in the main repository. Understanding this foundation is essential before selecting a specific runtime.

Python Reference Implementation

The voxcpm package defines the four-stage pipeline used by all downstream projects. In src/voxcpm/model/voxcpm.py, the top-level VoxCPM class orchestrates the encoder, language model, and decoder components.

Key source files that external tools import or replicate include:

The Python API serves as the reference implementation that the C++, ONNX, and Rust tools either wrap or re-implement.

VoxCPM.cpp: High-Performance C++ Runtime

VoxCPM.cpp provides GGML/GGUF-based inference for CPU, CUDA, or Vulkan backends. This implementation eliminates Python overhead, offering extremely low-memory, high-speed execution suitable for resource-constrained environments.

The tool parses model weights converted to GGUF format and runs the pipeline natively. Conversion scripts provided in the VoxCPM.cpp repository transform the original PyTorch checkpoints into the quantized GGUF format.


# Convert PyTorch checkpoint to GGUF format

voxcpm-convert --ckpt VoxCPM2.pt --gguf VoxCPM2.gguf

# Run inference on CPU

./voxcpm -m VoxCPM2.gguf -t "A calm, female narrator reads a story." -o story.wav

The C++ runtime maintains compatibility with the core model definition while removing the Python dependency entirely, making it ideal for production deployments and edge inference.

VoxCPM-ONNX: Cross-Platform Export and Inference

VoxCPM-ONNX enables export of the model to ONNX format and provides a lightweight ONNX Runtime wrapper for cross-platform CPU inference. This approach bridges the gap between the Python reference and fully native implementations.

The export process uses a script in the main repository (typically located at scripts/export_onnx.py) to trace the VoxCPM graph with torch.onnx.export:

python -m pip install onnx onnxruntime
python scripts/export_onnx.py \
    --model_dir ./pretrained_models/VoxCPM2 \
    --output voxcpm.onnx

Once exported, the ONNX model integrates with the VoxCPM-ONNX runtime wrapper:

import onnxruntime as ort
import numpy as np

sess = ort.InferenceSession("voxcpm.onnx")

# Prepare inputs according to the wrapper's API

audio = sess.run(None, {"text": np.array(["Hello from ONNX!"])})[0]

This toolchain is particularly valuable for Windows deployments and environments where Python installation is undesirable but full C++ compilation is impractical.

ComfyUI Extensions: Visual Workflow Integration

The VoxCPM ecosystem includes multiple node-based integrations for ComfyUI, allowing users to embed text-to-speech pipelines inside image-generation workflows without writing code.

ComfyUI-VoxCPM

The primary ComfyUI extension creates custom nodes that import voxcpm directly or load ONNX/GGUF models. Nodes expose parameters such as text, reference_wav, and control_prompt, forwarding data through the standard LocEnc → TSLM → RALM → LocDiT pipeline.

Installation requires placing the package inside the ComfyUI custom_nodes folder. Once installed, users can add a VoxCPM TTS node, configure the text field, and connect outputs to Audio Playback or Save Audio nodes.

{
  "type": "VoxCPM_TTS",
  "inputs": {
    "text": "(Cheerful, high-energy voice)Welcome to ComfyUI with VoxCPM!",
    "reference_wav": null
  }
}

ComfyUI-VoxCPMTTS

An alternative extension, ComfyUI-VoxCPMTTS, focuses exclusively on TTS-only nodes, providing a streamlined interface for audio generation without image-generation coupling. Both extensions allow seamless chaining with standard ComfyUI image nodes, enabling synchronized audiovisual content creation.

Additional Community Tools

Beyond the primary trio of VoxCPM.cpp, ONNX, and ComfyUI tools, the ecosystem includes specialized implementations:

  • VoxCPMANE – Targets the Apple Neural Engine (ANE) for on-device inference on iOS and macOS devices
  • voxcpm_rs – A Rust re-implementation offering high-performance, safe concurrency for systems programming contexts
  • TTS WebUI – A browser-based Gradio/Streamlit interface that calls the Python library directly, suitable for rapid prototyping

All these tools import, export, or re-implement the core model definition found in src/voxcpm/model/voxcpm.py, ensuring consistent behavior across runtimes.

Summary

  • VoxCPM.cpp delivers GGML/GGUF-based C++ inference for CPU, CUDA, and Vulkan, eliminating Python overhead while maintaining the LocEnc → TSLM → RALM → LocDiT pipeline
  • VoxCPM-ONNX provides cross-platform deployment via ONNX Runtime, using export scripts from the main repository to convert the PyTorch model
  • ComfyUI extensions (ComfyUI-VoxCPM and ComfyUI-VoxCPMTTS) enable node-based visual programming integration for TTS workflows
  • All ecosystem tools reference the canonical Python implementation in src/voxcpm/model/voxcpm.py and its constituent modules (local_encoder.py, local_dit.py, audio_vae_v2.py)
  • The ecosystem supports diverse deployment targets including edge devices (VoxCPMANE), systems programming environments (voxcpm_rs), and browser-based UIs

Frequently Asked Questions

What is the difference between VoxCPM.cpp and VoxCPM-ONNX?

VoxCPM.cpp is a native C++ implementation using the GGML/GGUF format for maximum performance and minimal memory footprint, requiring compilation from source. VoxCPM-ONNX exports the model to the standardized ONNX format and runs via ONNX Runtime, offering easier cross-platform deployment without compilation but potentially higher latency than the optimized C++ implementation.

Can I use the ComfyUI extensions without installing the full Python package?

No, the ComfyUI extensions typically require the voxcpm Python package to be installed in the environment, or they must load the exported ONNX/GGUF models. The nodes act as interfaces to the underlying inference engines rather than standalone implementations. Refer to the specific extension repository (ComfyUI-VoxCPM or ComfyUI-VoxCPMTTS) for exact dependency requirements.

Which ecosystem tool offers the lowest latency for CPU-only inference?

VoxCPM.cpp provides the lowest latency for CPU-only scenarios because it uses quantized GGUF weights and the highly optimized GGML inference engine written in C++. The Python reference implementation and ONNX Runtime wrappers introduce additional overhead from the Python interpreter and runtime abstraction layers, respectively.

Where is the AudioVAE V2 decoder implemented in the source code?

The AudioVAE V2 decoder is implemented in src/voxcpm/modules/audiovae/audio_vae_v2.py within the main OpenBMB/VoxCPM repository. This module handles the final super-resolution stage that produces 48 kHz audio output, and it is referenced by all ecosystem tools that maintain full fidelity to the original model architecture.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →