# Pocket TTS Alternative Implementations: ONNX, WebAssembly, MLX, and C++ Ports

> Explore Pocket TTS alternative implementations including ONNX Web, WebAssembly, MLX, and C++ ports. Deploy TTS in browsers, on Apple Silicon, and embedded systems.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: deep-dive
- Published: 2026-07-11

---

**Community-maintained ports of Pocket TTS exist for ONNX Runtime Web, WebAssembly, Apple's MLX framework, and native C++, enabling deployment across browsers, Apple Silicon, and embedded systems.**

Pocket TTS from kyutai-labs is a lightweight CPU-only text-to-speech model with approximately 100 million parameters. While the reference implementation ships as a pure Python package, the repository's clean architecture has spawned several alternative implementations targeting different runtimes and hardware platforms.

## Understanding the Core Architecture

### FlowLMModel and MimiModel Pipeline

The reference implementation in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) orchestrates a two-stage pipeline. First, the `FlowLMModel` class in [`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py) generates latent audio codes using a flow-based transformer. Second, the `MimiModel` class in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py) decodes these latents into raw audio via a neural audio codec.

### Streaming Design and State Management

The architecture relies on streaming transformer blocks with RoPE embeddings defined in [`pocket_tts/modules/transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/transformer.py) and stateful KV-caching implemented in [`pocket_tts/modules/stateful_module.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/stateful_module.py). This design enables low-latency voice cloning and efficient CPU inference that alternative implementations preserve across ports.

## ONNX and WebAssembly Implementations

### pocket-tts-onnx-export

The **pocket-tts-onnx-export** project exports the core Flow-LM model to an `.onnx` file for hardware-agnostic inference. This enables deployment via ONNX Runtime Web, allowing the model to execute in browsers using WebAssembly without server-side processing.

### wasm-pocket-tts

**wasm-pocket-tts** provides a Rust re-implementation that compiles to WebAssembly via the XN framework. This port delivers client-side TTS inference entirely within the browser, ideal for privacy-preserving applications that avoid network requests.

### sherpa-onnx

The **sherpa-onnx** project wraps the exported ONNX model and provides bindings for 12 programming languages plus WebAssembly support. This creates a comprehensive cross-platform solution for embedding Pocket TTS in diverse application stacks.

## Apple Silicon Optimization with MLX

### pocket-tts-mlx

The **pocket-tts-mlx** implementation rewrites the inference stack to use Apple's MLX library. This alternative implementation delivers accelerated CPU-only performance on M-series chips by leveraging Apple's optimized ML stack while maintaining compatibility with the original model weights.

## Native C++ Implementation

### PocketTTS.cpp

**PocketTTS.cpp** offers a single-file C++ runtime built on ONNX Runtime. This implementation provides a command-line interface, HTTP server, and FFI C API for embedding the TTS engine in native applications with minimal overhead. The port uses the same `.onnx` model exports as the WebAssembly implementations but targets native binaries.

## Additional Community Ports

### pocket-tts-csharp

The **pocket-tts-csharp** project provides a .NET library built on TorchSharp that mirrors the Python API. This enables C# developers to integrate Pocket TTS using familiar patterns while maintaining feature parity with the reference implementation.

## Code Examples

### Reference Python API

```python
from pocket_tts import TTSModel
import scipy.io.wavfile

# Load the model (downloads weights on first use)

tts = TTSModel.load_model()

# Create a voice state (can also pass a local wav file)

voice = tts.get_state_for_audio_prompt("alba")

# Generate speech from text

audio = tts.generate_audio(voice, "Hello, world! This is Pocket TTS.")
scipy.io.wavfile.write("hello.wav", tts.sample_rate, audio.numpy())

```

### Browser-based ONNX WebAssembly

```html
<script type="module">
import * as ort from "https://cdn.jsdelivr.net/npm/onnxruntime-web/dist/ort.min.js";

async function runTTS(text) {
  // Load the exported ONNX model (generated by pocket‑tts‑onnx‑export)
  const session = await ort.InferenceSession.create("pocket_tts.onnx");

  // Tokenize the text using the same SentencePiece vocab (load from HF)
  const tokens = await fetch("tokenizer.model").then(r => r.arrayBuffer())
                .then(model => sentencepieceEncode(model, text));

  // Prepare input tensors (batch=1, seq_len = tokens.length)
  const feeds = { "input_ids": new ort.Tensor("int64", tokens, [1, tokens.length]) };

  // Perform inference
  const results = await session.run(feeds);
  const audioLatents = results["latent_output"];   // shape: [1, frames, latent_dim]

  // Decode latents with the Mimi codec (WebAssembly implementation)
  const audio = await decodeWithMimi(audioLatents);
  playAudio(audio);
}
</script>

```

### C++ ONNX Runtime

```cpp
#include "pocket_tts_cpp.h"

int main() {
    // Initialise the runtime with the exported ONNX model
    PocketTTS tts("pocket_tts.onnx");

    // Load a voice embedding (pre‑computed .safetensors)
    tts.load_voice("alba.safetensors");

    // Generate audio for a given string
    std::vector<float> wav = tts.generate("Hello from C++!");

    // Write to a WAV file (using stb_vorbis or similar)
    write_wav("hello.wav", wav, 24000);
}

```

## Key Source Files for Porting

When building alternative implementations, these reference files from the kyutai-labs/pocket-tts repository provide the canonical architecture:

- **[`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py)**: Orchestrates the full TTS pipeline including voice caching and streaming generation
- **[`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py)**: Implements the Flow-LM transformer producing latent audio codes
- **[`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py)**: Wraps the Mimi neural audio codec for encoding and decoding
- **[`pocket_tts/modules/transformer.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/transformer.py)**: Core streaming transformer blocks with RoPE embeddings
- **[`pocket_tts/modules/stateful_module.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/modules/stateful_module.py)**: Base class providing KV-cache and state handling for streaming
- **[`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py)**: SentencePiece tokenizer and embedding lookup for text conditioning
- **[`pocket_tts/utils/weights_loading.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/weights_loading.py)**: Handles `hf://` URLs, downloading, and caching of model weights

## Summary

- **Alternative implementations** of Pocket TTS exist for ONNX Runtime Web, WebAssembly, Apple's MLX, and C++, enabling deployment across browsers, Apple Silicon, and embedded systems.
- All community ports preserve the **core architecture** from the reference Python implementation, including the Flow-LM transformer in [`pocket_tts/models/flow_lm.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/flow_lm.py) and the Mimi codec in [`pocket_tts/models/mimi.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/mimi.py).
- **ONNX Runtime** provides hardware-agnostic inference for both browser WebAssembly and native C++ applications via projects like **PocketTTS.cpp** and **sherpa-onnx**.
- **MLX optimization** delivers superior performance on Apple Silicon by rewriting the inference stack to use Apple's accelerated ML framework.
- The **streaming design**, voice-cloning capabilities, and low-latency characteristics remain consistent across all alternative implementations.

## Frequently Asked Questions

### What is the primary difference between the Python reference and the ONNX alternative implementations?

The Python reference implementation uses PyTorch for model execution, while ONNX alternative implementations export the trained weights to a hardware-agnostic graph format. This allows the model to run on ONNX Runtime across different languages and platforms, including WebAssembly in browsers and native C++ binaries, without requiring PyTorch dependencies.

### Can I use Pocket TTS in a browser without sending data to a server?

Yes. The **wasm-pocket-tts** and **pocket-tts-onnx-export** projects enable fully client-side inference via WebAssembly. These implementations compile the model to run in the browser, ensuring privacy by processing text and generating audio locally without network requests.

### Which alternative implementation should I choose for Apple Silicon Macs?

For optimal performance on Apple Silicon, use **pocket-tts-mlx**, which rewrites the inference stack to use Apple's MLX library. This implementation specifically targets M-series chips to deliver faster CPU-only performance compared to the standard PyTorch or ONNX alternatives.

### Are the alternative implementations compatible with the same voice files?

Yes. Community ports maintain compatibility with the original model weights and voice embeddings. The C++ implementation loads `.safetensors` voice files, while WebAssembly and ONNX versions use the same tokenizer vocabulary and model architecture defined in [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py), ensuring consistent voice cloning across all runtimes.