How VoxCPM's Tokenizer-Free Diffusion Autoregressive Architecture Works
VoxCPM generates speech by directly mapping text to audio through a character-level language model that drives a local diffusion transformer in an autoregressive loop, eliminating the need for traditional audio tokenizers.
VoxCPM (also referred to as VoxCPM2) is an open-source text-to-speech system developed by OpenBMB that implements a tokenizer-free diffusion autoregressive architecture. Unlike conventional TTS models that rely on discrete audio tokens from external vocoders or neural codecs, this architecture couples a causal language model with a continuous
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →