Pocket TTS CLI Commands: generate, serve, and export-voice Options Explained
The Pocket TTS CLI provides three Typer-based commands—generate, serve, and export-voice—that load models via TTSModel.load_model(), accept voice prompts from files or URLs, and output either WAV audio, an HTTP stream, or serialized voice states.
The kyutai-labs/pocket-tts repository ships a lightweight, self-contained text-to-speech toolkit built on the Flow-MIMI architecture. Its command-line interface—implemented in pocket_tts/main.py using Typer—exposes three primary subcommands that handle everything from single audio file generation to HTTP API serving and voice cloning state serialization.
The generate Command
The generate command synthesizes speech from text and streams the output to a WAV file. It supports reading text from STDIN (use -) and conditions the output on built-in voices, remote URLs, or local audio files via tts_model.get_state_for_audio_prompt().
Parameters for Speech Generation
The following options are defined in pocket_tts/main.py and control the synthesis pipeline:
--text– The text to synthesize. Pass-to read from STDIN.--voice– Path or URL to conditioning audio (voice to clone). Supports HuggingFace URLs (hf://...) and local files. If omitted, the built-in default for the selected language is used.--language– Language or model identifier (e.g.,english,italian_24l). Mutually exclusive with--config.--config– Path to a custom YAML configuration file. Mutually exclusive with--language.--output-path– Destination for the generated WAV file. Defaults to./tts_output.wav. Use-to stream raw audio to STDOUT.--device– Torch device for inference. Defaults to"cpu".--quantize– Apply 8-bit quantization to reduce RAM usage. Boolean flag.--lsd-decode-steps– Number of decoding steps for the Lagrangian Self-Distillation pipeline. Defaults toDEFAULT_LSD_DECODE_STEPSfrompocket_tts/default_parameters.py.--temperature– Sampling temperature for generation. Defaults toDEFAULT_TEMPERATURE.--noise-clamp– Maximum noise amplitude. Defaults toDEFAULT_NOISE_CLAMP.--eos-threshold– Threshold for end-of-speech detection. Defaults toDEFAULT_EOS_THRESHOLD.--frames-after-eos– Extra frames emitted after EOS is detected. Defaults toDEFAULT_FRAMES_AFTER_EOS.--max-tokens– Upper bound on tokens per generation chunk. Defaults toMAX_TOKEN_PER_CHUNK.--quiet(-q) – Suppress log output.
Generate Usage Examples
Generate a short clip using defaults:
uvx pocket-tts generate
Synthesize custom text with a specific voice and output path:
pocket-tts generate \
--text "Hello from Pocket TTS!" \
--voice "alba" \
--output-path ./hello.wav
Read text from STDIN and stream audio to another program:
echo "Streaming test" | pocket-tts generate --text - --output-path -
The serve Command
The serve command launches a FastAPI server that keeps the model warm in memory, streaming WAV responses via HTTP POST requests. This is defined in pocket_tts/main.py starting at line 86.
Server Parameters
--host– Hostname to bind the server. Defaults to"localhost".--port– TCP port. Defaults to8000.--reload– Enable auto-reload for development. Boolean flag.--language– Language selector (same values asgenerate).--config– Optional custom YAML config path.--quantize– Enable 8-bit quantization for the server process.
Launching the API
Start the server on all interfaces:
uvx pocket-tts serve --host 0.0.0.0 --port 8080
Once running, send POST requests to http://localhost:8080/tts with a JSON payload containing the text and optional voice parameters.
The export-voice Command
The export-voice command pre-processes an audio file into a serialized .safetensors state, enabling faster voice loading in subsequent generate calls. This avoids re-processing the conditioning audio on every synthesis run.
Voice Serialization Options
audio-path(positional, required) – Path or directory containing the source audio to convert.export-path(positional, required) – Destination file or directory for the.safetensorsoutput.--language– Language/model selector (same set asgenerate).--config– Optional custom YAML config.--quiet(-q) – Suppress log output.
Export Example
Convert a custom voice recording for fast reuse:
pocket-tts export-voice \
./my_voice.wav \
./my_voice.safetensors \
--language italian_24l
Shared Model Loading Architecture
All three commands rely on TTSModel.load_model() located in pocket_tts/models/tts_model.py to download weights and initialize the Flow-LM and Mimi codec. Voice conditioning is handled through get_state_for_audio_prompt(), which converts audio inputs into model states regardless of whether the source is a built-in voice, a URL, or a local file.
The CLI enforces mutual exclusivity between --language and --config at the argument parsing level, ensuring users either select a predefined model configuration or provide a custom YAML file, never both.
Summary
generatestreams synthesized audio to a WAV file with fine-grained control over sampling temperature, decode steps, and device placement.servelaunches a FastAPI server for HTTP-based TTS requests, supporting the same model configurations as the CLI but keeping the model resident in memory.export-voiceconverts raw audio prompts into.safetensorsfiles for rapid voice reloading without repeated preprocessing.- All commands share the same
TTSModelinitialization path and support 8-bit quantization via the--quantizeflag to reduce memory consumption.
Frequently Asked Questions
Can I use a custom model configuration instead of the built-in languages?
Yes. Pass --config /path/to/config.yaml to any command. This option is mutually exclusive with --language; you must use one or the other, not both. The YAML file should define the model architecture and checkpoint paths compatible with TTSModel.load_model().
How do I clone a voice from a custom audio file?
Use the --voice parameter in the generate command with a local file path (e.g., --voice ./sample.wav) or a remote URL. For frequent reuse, first run export-voice on the audio to create a .safetensors file, then pass that file path to --voice for faster loading.
What is the difference between the default voices and exported voice files?
Built-in voices (referenced by name like "alba") are downloaded automatically from predefined origins managed in pocket_tts/utils/utils.py. Exported voice files (.safetensors) contain pre-computed model states created by export-voice, eliminating the need to reprocess the conditioning audio on every generation call.
Does the serve command support GPU acceleration?
The serve command accepts the --quantize flag for 8-bit quantization to reduce memory usage, but it does not expose a --device parameter in the current implementation. For GPU inference, use the generate command with --device cuda, as the server defaults to CPU-based execution according to the parameter definitions in pocket_tts/main.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →