# Understanding the VoxCPM Pipeline: LocEnc → TTSLM → RALM → LocDiT

> Explore the VoxCPM pipeline: LocEnc, TTSLM, RALM and LocDiT. Understand how these modules process speech and audio for advanced text-to-speech generation and refinement via OpenBMB.

- Repository: [OpenBMB/VoxCPM](https://github.com/OpenBMB/VoxCPM)
- Tags: deep-dive
- Published: 2026-04-10

---

**The VoxCPM pipeline processes speech through four specialized neural modules: a Local Encoder (LocEnc) converts acoustic features into embeddings, a Text-to-Speech Language Model (TTSLM) jointly processes text and audio representations, a Residual Acoustic LM (RALM) refines acoustic details, and a Local DiT (LocDiT) diffuses the combined hidden states into final mel-spectrograms.**

The OpenBMB/VoxCPM repository implements this novel four-stage architecture for high-f