How AudioVAE V2 Achieves Native 48kHz Speech Output in OpenBMB/VoxCPM

AudioVAE V2 generates native 48kHz speech by conditioning the decoder on the target sample rate through learned bucket-specific transformations, eliminating the need for external resampling libraries.

The OpenBMB/VoxCPM repository implements AudioVAE V2 as a neural audio codec that encodes 16kHz speech into a compressed latent representation while decoding directly to 48kHz. This architecture achieves native 48kHz speech output through internal sample-rate conditioning mechanisms that adapt the decoder's spectral characteristics to match high-frequency

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →