Pocket TTS Alternative Implementations: ONNX, WebAssembly, MLX, and C++ Ports

Community-maintained ports of Pocket TTS exist for ONNX Runtime Web, WebAssembly, Apple's MLX framework, and native C++, enabling deployment across browsers, Apple Silicon, and embedded systems.

Pocket TTS from kyutai-labs is a lightweight CPU-only text-to-speech model with approximately 100 million parameters. While the reference implementation ships as a pure Python package, the repository's clean architecture has spawned several alternative implementations targeting different runtimes and hardware platforms.

Understanding the Core Architecture

FlowLMModel and MimiModel Pipeline

The reference implementation in pocket_tts/models/tts_model.py orchestrates a two-stage pipeline. First, the FlowLMModel class in pocket_tts/models/flow_lm.py generates latent audio codes using a flow-based transformer. Second, the MimiModel class in pocket_tts/models/mimi.py decodes these latents into raw audio via a neural audio codec.

Streaming Design and State Management

The architecture relies on streaming transformer blocks with RoPE embeddings defined in pocket_tts/modules/transformer.py and stateful KV-caching implemented in pocket_tts/modules/stateful_module.py. This design enables low-latency voice cloning and efficient CPU inference that alternative implementations preserve across ports.

ONNX and WebAssembly Implementations

pocket-tts-onnx-export

The pocket-tts-onnx-export project exports the core Flow-LM model to an .onnx file for hardware-agnostic inference. This enables deployment via ONNX Runtime Web, allowing the model to execute in browsers using WebAssembly without server-side processing.

wasm-pocket-tts

wasm-pocket-tts provides a Rust re-implementation that compiles to WebAssembly via the XN framework. This port delivers client-side TTS inference entirely within the browser, ideal for privacy-preserving applications that avoid network requests.

sherpa-onnx

The sherpa-onnx project wraps the exported ONNX model and provides bindings for 12 programming languages plus WebAssembly support. This creates a comprehensive cross-platform solution for embedding Pocket TTS in diverse application stacks.

Apple Silicon Optimization with MLX

pocket-tts-mlx

The pocket-tts-mlx implementation rewrites the inference stack to use Apple's MLX library. This alternative implementation delivers accelerated CPU-only performance on M-series chips by leveraging Apple's optimized ML stack while maintaining compatibility with the original model weights.

Native C++ Implementation

PocketTTS.cpp

PocketTTS.cpp offers a single-file C++ runtime built on ONNX Runtime. This implementation provides a command-line interface, HTTP server, and FFI C API for embedding the TTS engine in native applications with minimal overhead. The port uses the same .onnx model exports as the WebAssembly implementations but targets native binaries.

Additional Community Ports

pocket-tts-csharp

The pocket-tts-csharp project provides a .NET library built on TorchSharp that mirrors the Python API. This enables C# developers to integrate Pocket TTS using familiar patterns while maintaining feature parity with the reference implementation.

Code Examples

Reference Python API

from pocket_tts import TTSModel
import scipy.io.wavfile

# Load the model (downloads weights on first use)

tts = TTSModel.load_model()

# Create a voice state (can also pass a local wav file)

voice = tts.get_state_for_audio_prompt("alba")

# Generate speech from text

audio = tts.generate_audio(voice, "Hello, world! This is Pocket TTS.")
scipy.io.wavfile.write("hello.wav", tts.sample_rate, audio.numpy())

Browser-based ONNX WebAssembly

<script type="module">
import * as ort from "https://cdn.jsdelivr.net/npm/onnxruntime-web/dist/ort.min.js";

async function runTTS(text) {
  // Load the exported ONNX model (generated by pocket‑tts‑onnx‑export)
  const session = await ort.InferenceSession.create("pocket_tts.onnx");

  // Tokenize the text using the same SentencePiece vocab (load from HF)
  const tokens = await fetch("tokenizer.model").then(r => r.arrayBuffer())
                .then(model => sentencepieceEncode(model, text));

  // Prepare input tensors (batch=1, seq_len = tokens.length)
  const feeds = { "input_ids": new ort.Tensor("int64", tokens, [1, tokens.length]) };

  // Perform inference
  const results = await session.run(feeds);
  const audioLatents = results["latent_output"];   // shape: [1, frames, latent_dim]

  // Decode latents with the Mimi codec (WebAssembly implementation)
  const audio = await decodeWithMimi(audioLatents);
  playAudio(audio);
}
</script>

C++ ONNX Runtime

#include "pocket_tts_cpp.h"

int main() {
    // Initialise the runtime with the exported ONNX model
    PocketTTS tts("pocket_tts.onnx");

    // Load a voice embedding (pre‑computed .safetensors)
    tts.load_voice("alba.safetensors");

    // Generate audio for a given string
    std::vector<float> wav = tts.generate("Hello from C++!");

    // Write to a WAV file (using stb_vorbis or similar)
    write_wav("hello.wav", wav, 24000);
}

Key Source Files for Porting

When building alternative implementations, these reference files from the kyutai-labs/pocket-tts repository provide the canonical architecture:

Summary

  • Alternative implementations of Pocket TTS exist for ONNX Runtime Web, WebAssembly, Apple's MLX, and C++, enabling deployment across browsers, Apple Silicon, and embedded systems.
  • All community ports preserve the core architecture from the reference Python implementation, including the Flow-LM transformer in pocket_tts/models/flow_lm.py and the Mimi codec in pocket_tts/models/mimi.py.
  • ONNX Runtime provides hardware-agnostic inference for both browser WebAssembly and native C++ applications via projects like PocketTTS.cpp and sherpa-onnx.
  • MLX optimization delivers superior performance on Apple Silicon by rewriting the inference stack to use Apple's accelerated ML framework.
  • The streaming design, voice-cloning capabilities, and low-latency characteristics remain consistent across all alternative implementations.

Frequently Asked Questions

What is the primary difference between the Python reference and the ONNX alternative implementations?

The Python reference implementation uses PyTorch for model execution, while ONNX alternative implementations export the trained weights to a hardware-agnostic graph format. This allows the model to run on ONNX Runtime across different languages and platforms, including WebAssembly in browsers and native C++ binaries, without requiring PyTorch dependencies.

Can I use Pocket TTS in a browser without sending data to a server?

Yes. The wasm-pocket-tts and pocket-tts-onnx-export projects enable fully client-side inference via WebAssembly. These implementations compile the model to run in the browser, ensuring privacy by processing text and generating audio locally without network requests.

Which alternative implementation should I choose for Apple Silicon Macs?

For optimal performance on Apple Silicon, use pocket-tts-mlx, which rewrites the inference stack to use Apple's MLX library. This implementation specifically targets M-series chips to deliver faster CPU-only performance compared to the standard PyTorch or ONNX alternatives.

Are the alternative implementations compatible with the same voice files?

Yes. Community ports maintain compatibility with the original model weights and voice embeddings. The C++ implementation loads .safetensors voice files, while WebAssembly and ONNX versions use the same tokenizer vocabulary and model architecture defined in pocket_tts/conditioners/text.py, ensuring consistent voice cloning across all runtimes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →