Pocket TTS Alternative Implementations: ONNX, WebAssembly, MLX, and C++ Ports
Community-maintained ports of Pocket TTS exist for ONNX Runtime Web, WebAssembly, Apple's MLX framework, and native C++, enabling deployment across browsers, Apple Silicon, and embedded systems.
Pocket TTS from kyutai-labs is a lightweight CPU-only text-to-speech model with approximately 100 million parameters. While the reference implementation ships as a pure Python package, the repository's clean architecture has spawned several alternative implementations targeting different runtimes and hardware platforms.
Understanding the Core Architecture
FlowLMModel and MimiModel Pipeline
The reference implementation in pocket_tts/models/tts_model.py orchestrates a two-stage pipeline. First, the FlowLMModel class in pocket_tts/models/flow_lm.py generates latent audio codes using a flow-based transformer. Second, the MimiModel class in pocket_tts/models/mimi.py decodes these latents into raw audio via a neural audio codec.
Streaming Design and State Management
The architecture relies on streaming transformer blocks with RoPE embeddings defined in pocket_tts/modules/transformer.py and stateful KV-caching implemented in pocket_tts/modules/stateful_module.py. This design enables low-latency voice cloning and efficient CPU inference that alternative implementations preserve across ports.
ONNX and WebAssembly Implementations
pocket-tts-onnx-export
The pocket-tts-onnx-export project exports the core Flow-LM model to an .onnx file for hardware-agnostic inference. This enables deployment via ONNX Runtime Web, allowing the model to execute in browsers using WebAssembly without server-side processing.
wasm-pocket-tts
wasm-pocket-tts provides a Rust re-implementation that compiles to WebAssembly via the XN framework. This port delivers client-side TTS inference entirely within the browser, ideal for privacy-preserving applications that avoid network requests.
sherpa-onnx
The sherpa-onnx project wraps the exported ONNX model and provides bindings for 12 programming languages plus WebAssembly support. This creates a comprehensive cross-platform solution for embedding Pocket TTS in diverse application stacks.
Apple Silicon Optimization with MLX
pocket-tts-mlx
The pocket-tts-mlx implementation rewrites the inference stack to use Apple's MLX library. This alternative implementation delivers accelerated CPU-only performance on M-series chips by leveraging Apple's optimized ML stack while maintaining compatibility with the original model weights.
Native C++ Implementation
PocketTTS.cpp
PocketTTS.cpp offers a single-file C++ runtime built on ONNX Runtime. This implementation provides a command-line interface, HTTP server, and FFI C API for embedding the TTS engine in native applications with minimal overhead. The port uses the same .onnx model exports as the WebAssembly implementations but targets native binaries.
Additional Community Ports
pocket-tts-csharp
The pocket-tts-csharp project provides a .NET library built on TorchSharp that mirrors the Python API. This enables C# developers to integrate Pocket TTS using familiar patterns while maintaining feature parity with the reference implementation.
Code Examples
Reference Python API
from pocket_tts import TTSModel
import scipy.io.wavfile
# Load the model (downloads weights on first use)
tts = TTSModel.load_model()
# Create a voice state (can also pass a local wav file)
voice = tts.get_state_for_audio_prompt("alba")
# Generate speech from text
audio = tts.generate_audio(voice, "Hello, world! This is Pocket TTS.")
scipy.io.wavfile.write("hello.wav", tts.sample_rate, audio.numpy())
Browser-based ONNX WebAssembly
<script type="module">
import * as ort from "https://cdn.jsdelivr.net/npm/onnxruntime-web/dist/ort.min.js";
async function runTTS(text) {
// Load the exported ONNX model (generated by pocket‑tts‑onnx‑export)
const session = await ort.InferenceSession.create("pocket_tts.onnx");
// Tokenize the text using the same SentencePiece vocab (load from HF)
const tokens = await fetch("tokenizer.model").then(r => r.arrayBuffer())
.then(model => sentencepieceEncode(model, text));
// Prepare input tensors (batch=1, seq_len = tokens.length)
const feeds = { "input_ids": new ort.Tensor("int64", tokens, [1, tokens.length]) };
// Perform inference
const results = await session.run(feeds);
const audioLatents = results["latent_output"]; // shape: [1, frames, latent_dim]
// Decode latents with the Mimi codec (WebAssembly implementation)
const audio = await decodeWithMimi(audioLatents);
playAudio(audio);
}
</script>
C++ ONNX Runtime
#include "pocket_tts_cpp.h"
int main() {
// Initialise the runtime with the exported ONNX model
PocketTTS tts("pocket_tts.onnx");
// Load a voice embedding (pre‑computed .safetensors)
tts.load_voice("alba.safetensors");
// Generate audio for a given string
std::vector<float> wav = tts.generate("Hello from C++!");
// Write to a WAV file (using stb_vorbis or similar)
write_wav("hello.wav", wav, 24000);
}
Key Source Files for Porting
When building alternative implementations, these reference files from the kyutai-labs/pocket-tts repository provide the canonical architecture:
pocket_tts/models/tts_model.py: Orchestrates the full TTS pipeline including voice caching and streaming generationpocket_tts/models/flow_lm.py: Implements the Flow-LM transformer producing latent audio codespocket_tts/models/mimi.py: Wraps the Mimi neural audio codec for encoding and decodingpocket_tts/modules/transformer.py: Core streaming transformer blocks with RoPE embeddingspocket_tts/modules/stateful_module.py: Base class providing KV-cache and state handling for streamingpocket_tts/conditioners/text.py: SentencePiece tokenizer and embedding lookup for text conditioningpocket_tts/utils/weights_loading.py: Handleshf://URLs, downloading, and caching of model weights
Summary
- Alternative implementations of Pocket TTS exist for ONNX Runtime Web, WebAssembly, Apple's MLX, and C++, enabling deployment across browsers, Apple Silicon, and embedded systems.
- All community ports preserve the core architecture from the reference Python implementation, including the Flow-LM transformer in
pocket_tts/models/flow_lm.pyand the Mimi codec inpocket_tts/models/mimi.py. - ONNX Runtime provides hardware-agnostic inference for both browser WebAssembly and native C++ applications via projects like PocketTTS.cpp and sherpa-onnx.
- MLX optimization delivers superior performance on Apple Silicon by rewriting the inference stack to use Apple's accelerated ML framework.
- The streaming design, voice-cloning capabilities, and low-latency characteristics remain consistent across all alternative implementations.
Frequently Asked Questions
What is the primary difference between the Python reference and the ONNX alternative implementations?
The Python reference implementation uses PyTorch for model execution, while ONNX alternative implementations export the trained weights to a hardware-agnostic graph format. This allows the model to run on ONNX Runtime across different languages and platforms, including WebAssembly in browsers and native C++ binaries, without requiring PyTorch dependencies.
Can I use Pocket TTS in a browser without sending data to a server?
Yes. The wasm-pocket-tts and pocket-tts-onnx-export projects enable fully client-side inference via WebAssembly. These implementations compile the model to run in the browser, ensuring privacy by processing text and generating audio locally without network requests.
Which alternative implementation should I choose for Apple Silicon Macs?
For optimal performance on Apple Silicon, use pocket-tts-mlx, which rewrites the inference stack to use Apple's MLX library. This implementation specifically targets M-series chips to deliver faster CPU-only performance compared to the standard PyTorch or ONNX alternatives.
Are the alternative implementations compatible with the same voice files?
Yes. Community ports maintain compatibility with the original model weights and voice embeddings. The C++ implementation loads .safetensors voice files, while WebAssembly and ONNX versions use the same tokenizer vocabulary and model architecture defined in pocket_tts/conditioners/text.py, ensuring consistent voice cloning across all runtimes.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →