# How VoxCPM's Tokenizer-Free Diffusion Autoregressive Architecture Works

> Discover how VoxCPM's tokenizer free diffusion autoregressive architecture generates speech directly from text to audio using a character level language model and local diffusion transformer.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: internals
- Published: 2026-04-10

---

**VoxCPM generates speech by directly mapping text to audio through a character-level language model that drives a local diffusion transformer in an autoregressive loop, eliminating the need for traditional audio tokenizers.**

VoxCPM (also referred to as VoxCPM2) is an open-source text-to-speech system developed by OpenBMB that implements a **tokenizer-free diffusion autoregressive architecture**. Unlike conventional TTS models that rely on discrete audio tokens from external vocoders or neural codecs, this architecture couples a causal language model with a continuous