How to Use the pocket-tts CLI for Speech Generation with Custom Options

The pocket-tts CLI provides a Typer-based command-line interface in pocket_tts/main.py that supports custom voices, generation parameters, and device selection for high-quality text-to-speech synthesis.

The pocket-tts command-line tool from the kyutai-labs/pocket-tts repository offers fine-grained control over speech generation through a comprehensive set of options. Built on the Typer framework, the CLI exposes the generate sub-command that orchestrates model loading, voice conditioning, and streaming audio output. This guide explains how to leverage custom options for tailoring voice characteristics, generation quality, and hardware utilization.

CLI Architecture and Entry Point

The CLI entry point resides in pocket_tts/main.py, where the cli_app.command() decorator registers the generate sub-command at lines 22-84. Typer automatically generates help text and parses arguments, mapping command-line flags directly to the underlying TTSModel methods.

When you execute pocket-tts generate, the tool chains together several discrete operations:

  • Logging configuration via enable_logging, which respects the --quiet flag to suppress informational output
  • Default text resolution using get_default_text_for_language when --text is omitted
  • Model initialization through TTSModel.load_model with optional INT8 quantization
  • Voice state preparation via TTSModel.get_state_for_audio_prompt with LRU caching for repeated prompts
  • Streaming generation using generate_audio_stream and the Mimi codec decoder
  • Audio output through stream_audio_chunks to either a file path or standard output

Core Generation Pipeline

Understanding the internal pipeline helps optimize your CLI usage with custom options.

Logging and Default Text Handling

By default, the CLI enables informational logging. Use --quiet to disable verbose output. If you omit the --text argument, the system calls get_default_text_for_language from pocket_tts/default_parameters.py to retrieve a language-specific demonstration prompt.

Model Loading and Quantization

The TTSModel.load_model function initializes the pre-trained weights and can apply INT8 quantization when you pass the --quantize flag. This reduces memory footprint and accelerates CPU inference without requiring GPU acceleration.

Voice Conditioning and Caching

Voice customization happens through two pathways:

  1. Default voices: If --voice is omitted, get_default_voice_for_language selects a built-in voice (such as alba for English) from the default parameters file
  2. Custom voices: You may specify local .wav files or HuggingFace URLs via --voice

The TTSModel.get_state_for_audio_prompt method extracts a model state encoding the speaker's timbre. The CLI caches these states using Python's lru_cache mechanism, ensuring efficient reuse when processing multiple texts with the same voice.

Streaming Audio Output

The generate_audio_stream generator yields audio chunks while a background thread decodes them through the Mimi codec. The stream_audio_chunks function handles the final write operation, directing output to either a specified file path (--output-path) or standard output when set to -.

Customization Options and Examples

The pocket-tts CLI exposes parameters that map directly to the generation routine, allowing precise control over temperature, decoding steps, noise clamping, and device placement.

Basic Usage and Language Selection

Generate speech using English defaults:

pocket-tts generate

Specify custom text and language:

pocket-tts generate --text "The quick brown fox jumps over the lazy dog."
pocket-tts generate --language italian --text "Ciao mondo"

Voice Customization

Provide a local voice sample or reference a HuggingFace-hosted voice:


# Local voice file

pocket-tts generate \
    --text "Testing custom voice conditioning." \
    --voice ./my_voice.wav

# HuggingFace URL

pocket-tts generate \
    --text "Bonjour, je parle avec une voix différente." \
    --voice hf://kyutai/tts-voices/alba-mackenna/casual.wav

Generation Quality Parameters

Control the stochasticity and depth of the generation process:

pocket-tts generate \
    --text "Testing higher temperature." \
    --temperature 1.2 \
    --lsd-decode-steps 3 \
    --noise-clamp 0.8

Tune end-of-speech detection to control truncation behavior:

pocket-tts generate \
    --text "Short phrase." \
    --eos-threshold -3.0 \
    --frames-after-eos 5

Limit token consumption per chunk (default is 50):

pocket-tts generate \
    --text "A long paragraph that will be split into multiple chunks." \
    --max-tokens 80

Output and Device Control

Target specific hardware and output destinations:


# Save to specific path

pocket-tts generate \
    --text "Saving to a custom location." \
    --output-path /tmp/output.wav

# Stream to stdout for piping

pocket-tts generate --text "Direct output" --output-path -

# GPU acceleration (defaults to CPU if unavailable)

pocket-tts generate --text "Running on GPU." --device cuda

# INT8 quantization for faster CPU inference

pocket-tts generate --text "Quantised model." --quantize

Summary

  • The pocket-tts CLI in pocket_tts/main.py uses Typer to expose the generate command with extensive customization options
  • Voice conditioning supports local files, HuggingFace URLs, or built-in defaults via get_default_voice_for_language
  • Model loading through TTSModel.load_model supports INT8 quantization via the --quantize flag
  • Generation parameters include --temperature, --lsd-decode-steps, --noise-clamp, --eos-threshold, and --max-tokens for fine-tuning output quality
  • Hardware control allows device selection via --device (cuda/cpu) and output redirection via --output-path
  • The streaming pipeline uses generate_audio_stream and the Mimi codec for efficient audio synthesis

Frequently Asked Questions

How do I use a custom voice file with the pocket-tts CLI?

Pass the path to your .wav file using the --voice argument. The CLI supports local file paths (e.g., --voice ./my_voice.wav) and HuggingFace URLs (e.g., --voice hf://kyutai/tts-voices/alba-mackenna/casual.wav). The TTSModel.get_state_for_audio_prompt function processes these files and caches the resulting voice state for efficient reuse.

What parameters control the quality and speed of speech generation?

Quality is controlled by --temperature (stochasticity, default varies), --lsd-decode-steps (decoding iterations), and --noise-clamp (noise limitation). For faster inference on CPU, enable INT8 quantization with the --quantize flag. Use --max-tokens to adjust the token budget per chunk, which affects both quality and generation speed.

Can I stream the audio output to another program instead of saving to a file?

Yes. Set --output-path to - (hyphen) to stream raw audio chunks to standard output. This allows you to pipe the output to other command-line tools or media players. The stream_audio_chunks function handles both file writes and stdout streaming.

How does the CLI handle language selection and default text?

When you specify --language (e.g., italian or english), the CLI calls get_default_text_for_language from pocket_tts/default_parameters.py to retrieve an appropriate demonstration text if you omit the --text argument. It also selects a default voice for that language using get_default_voice_for_language, falling back to built-in voices like alba for English.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →