# Pocket TTS CLI Commands: generate, serve, and export-voice Options Explained

> Master Pocket TTS CLI commands like generate serve and export_voice to create WAV audio stream audio or export voice states Learn to load models and process voice prompts efficiently

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: cli-tool
- Published: 2026-07-11

---

**The Pocket TTS CLI provides three Typer-based commands—`generate`, `serve`, and `export-voice`—that load models via `TTSModel.load_model()`, accept voice prompts from files or URLs, and output either WAV audio, an HTTP stream, or serialized voice states.**

The `kyutai-labs/pocket-tts` repository ships a lightweight, self-contained text-to-speech toolkit built on the Flow-MIMI architecture. Its command-line interface—implemented in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py) using Typer—exposes three primary subcommands that handle everything from single audio file generation to HTTP API serving and voice cloning state serialization.

## The `generate` Command

The `generate` command synthesizes speech from text and streams the output to a WAV file. It supports reading text from STDIN (use `-`) and conditions the output on built-in voices, remote URLs, or local audio files via `tts_model.get_state_for_audio_prompt()`.

### Parameters for Speech Generation

The following options are defined in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py) and control the synthesis pipeline:

- **`--text`** – The text to synthesize. Pass `-` to read from STDIN.
- **`--voice`** – Path or URL to conditioning audio (voice to clone). Supports HuggingFace URLs (`hf://...`) and local files. If omitted, the built-in default for the selected language is used.
- **`--language`** – Language or model identifier (e.g., `english`, `italian_24l`). **Mutually exclusive with `--config`**.
- **`--config`** – Path to a custom YAML configuration file. **Mutually exclusive with `--language`**.
- **`--output-path`** – Destination for the generated WAV file. Defaults to `./tts_output.wav`. Use `-` to stream raw audio to STDOUT.
- **`--device`** – Torch device for inference. Defaults to `"cpu"`.
- **`--quantize`** – Apply 8-bit quantization to reduce RAM usage. Boolean flag.
- **`--lsd-decode-steps`** – Number of decoding steps for the Lagrangian Self-Distillation pipeline. Defaults to `DEFAULT_LSD_DECODE_STEPS` from [`pocket_tts/default_parameters.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/default_parameters.py).
- **`--temperature`** – Sampling temperature for generation. Defaults to `DEFAULT_TEMPERATURE`.
- **`--noise-clamp`** – Maximum noise amplitude. Defaults to `DEFAULT_NOISE_CLAMP`.
- **`--eos-threshold`** – Threshold for end-of-speech detection. Defaults to `DEFAULT_EOS_THRESHOLD`.
- **`--frames-after-eos`** – Extra frames emitted after EOS is detected. Defaults to `DEFAULT_FRAMES_AFTER_EOS`.
- **`--max-tokens`** – Upper bound on tokens per generation chunk. Defaults to `MAX_TOKEN_PER_CHUNK`.
- **`--quiet` (`-q`)** – Suppress log output.

### Generate Usage Examples

Generate a short clip using defaults:

```bash
uvx pocket-tts generate

```

Synthesize custom text with a specific voice and output path:

```bash
pocket-tts generate \
    --text "Hello from Pocket TTS!" \
    --voice "alba" \
    --output-path ./hello.wav

```

Read text from STDIN and stream audio to another program:

```bash
echo "Streaming test" | pocket-tts generate --text - --output-path -

```

## The `serve` Command

The `serve` command launches a FastAPI server that keeps the model warm in memory, streaming WAV responses via HTTP POST requests. This is defined in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py) starting at line 86.

### Server Parameters

- **`--host`** – Hostname to bind the server. Defaults to `"localhost"`.
- **`--port`** – TCP port. Defaults to `8000`.
- **`--reload`** – Enable auto-reload for development. Boolean flag.
- **`--language`** – Language selector (same values as `generate`).
- **`--config`** – Optional custom YAML config path.
- **`--quantize`** – Enable 8-bit quantization for the server process.

### Launching the API

Start the server on all interfaces:

```bash
uvx pocket-tts serve --host 0.0.0.0 --port 8080

```

Once running, send POST requests to `http://localhost:8080/tts` with a JSON payload containing the text and optional voice parameters.

## The `export-voice` Command

The `export-voice` command pre-processes an audio file into a serialized `.safetensors` state, enabling faster voice loading in subsequent `generate` calls. This avoids re-processing the conditioning audio on every synthesis run.

### Voice Serialization Options

- **`audio-path`** (positional, required) – Path or directory containing the source audio to convert.
- **`export-path`** (positional, required) – Destination file or directory for the `.safetensors` output.
- **`--language`** – Language/model selector (same set as `generate`).
- **`--config`** – Optional custom YAML config.
- **`--quiet` (`-q`)** – Suppress log output.

### Export Example

Convert a custom voice recording for fast reuse:

```bash
pocket-tts export-voice \
    ./my_voice.wav \
    ./my_voice.safetensors \
    --language italian_24l

```

## Shared Model Loading Architecture

All three commands rely on `TTSModel.load_model()` located in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) to download weights and initialize the Flow-LM and Mimi codec. Voice conditioning is handled through `get_state_for_audio_prompt()`, which converts audio inputs into model states regardless of whether the source is a built-in voice, a URL, or a local file.

The CLI enforces mutual exclusivity between `--language` and `--config` at the argument parsing level, ensuring users either select a predefined model configuration or provide a custom YAML file, never both.

## Summary

- **`generate`** streams synthesized audio to a WAV file with fine-grained control over sampling temperature, decode steps, and device placement.
- **`serve`** launches a FastAPI server for HTTP-based TTS requests, supporting the same model configurations as the CLI but keeping the model resident in memory.
- **`export-voice`** converts raw audio prompts into `.safetensors` files for rapid voice reloading without repeated preprocessing.
- All commands share the same `TTSModel` initialization path and support 8-bit quantization via the `--quantize` flag to reduce memory consumption.

## Frequently Asked Questions

### Can I use a custom model configuration instead of the built-in languages?

Yes. Pass `--config /path/to/config.yaml` to any command. This option is mutually exclusive with `--language`; you must use one or the other, not both. The YAML file should define the model architecture and checkpoint paths compatible with `TTSModel.load_model()`.

### How do I clone a voice from a custom audio file?

Use the `--voice` parameter in the `generate` command with a local file path (e.g., `--voice ./sample.wav`) or a remote URL. For frequent reuse, first run `export-voice` on the audio to create a `.safetensors` file, then pass that file path to `--voice` for faster loading.

### What is the difference between the default voices and exported voice files?

Built-in voices (referenced by name like `"alba"`) are downloaded automatically from predefined origins managed in [`pocket_tts/utils/utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/utils.py). Exported voice files (`.safetensors`) contain pre-computed model states created by `export-voice`, eliminating the need to reprocess the conditioning audio on every generation call.

### Does the `serve` command support GPU acceleration?

The `serve` command accepts the `--quantize` flag for 8-bit quantization to reduce memory usage, but it does not expose a `--device` parameter in the current implementation. For GPU inference, use the `generate` command with `--device cuda`, as the server defaults to CPU-based execution according to the parameter definitions in [`pocket_tts/main.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/main.py).