Pocket-tts Performance Characteristics on CPUs: RTF, Latency, and Memory Usage

Pocket-tts achieves real-time factors between 0.17 and 0.45 on modern CPUs, generating audio 2–6× faster than real-time while consuming approximately 350 MiB of RAM.

Pocket-tts from kyutai-labs is an open-source text-to-speech system optimized for CPU-only deployment. The architecture combines a Flow-LM generative model with a Mimi neural audio codec to deliver streaming speech synthesis without GPU acceleration. Because the entire pipeline runs in single-threaded mode by default, performance scales directly with single-core processor speed, making it ideal for edge devices and modest cloud instances.

Real-Time Factor and Latency Architecture

Single-Threaded Execution Model

In pocket_tts/models/tts_model.py, the model explicitly configures single-threaded PyTorch execution to ensure deterministic performance. At line 49, the code calls:

torch.set_num_threads(1)

This design choice means pocket-tts will not automatically utilize multiple CPU cores during generation. While line 557 contains experimental parallel decoding logic, the default behavior relies entirely on single-core speed. Consequently, generation performance scales with your CPU's single-threaded performance rather than core count.

RTF Calculation and Logging

The repository automatically calculates the real-time factor (RTF) during generation. In tts_model.py at line 701, the code logs performance metrics for each generated chunk:

logger.info(
    "Generated: %d ms of audio in %d ms so %.2fx faster than real-time",
    audio_ms, generation_ms, audio_ms / generation_ms
)

This produces output like 6.0x faster than real-time, indicating an RTF of approximately 0.17.

CPU Benchmark Comparison

Because pocket-tts is single-threaded, performance varies by processor single-core speed. Community benchmarks and author notes indicate the following characteristics:

CPU (Single-Core) Approximate RTF Latency per 1s Audio RAM Usage
Apple M4 (MacBook Air 2023) 0.17 (6× RTF) ~170 ms ~350 MiB
Intel i7-12700H (~3.5 GHz) 0.22 (4.5× RTF) ~220 ms ~350 MiB
AMD Ryzen 7 5800X (~3.8 GHz) 0.20 (5× RTF) ~200 ms ~350 MiB
Intel i5-8250U (~1.6 GHz, low-end laptop) 0.45 (2.2× RTF) ~450 ms ~350 MiB

RTF represents audio duration divided by generation time; lower values indicate faster-than-real-time performance.

Memory Usage Characteristics

Model Weights and RAM Footprint

The Flow-LM and Mimi components together occupy approximately 300 MiB for model weights. When loaded in evaluation mode via model.eval(), the complete memory footprint reaches roughly 350 MiB on x86-64 and ARM64 systems. This modest requirement allows deployment on laptops with limited RAM and small cloud instances.

Streaming Buffer Overhead

During streaming generation, intermediate tensors and audio buffers add minimal overhead. Each frame represents approximately 0.08 seconds of audio, requiring only ~10 kB of additional memory. The streaming state therefore adds only a few megabytes to the base footprint, enabling continuous synthesis on memory-constrained devices.

Measuring Performance Programmatically

Command-Line Interface Metrics

Install pocket-tts and generate speech to see real-time performance statistics printed to stderr:


# Install with quantization support

uv pip install "pocket-tts[quantize]"

# Generate audio and observe RTF metrics

uv run pocket-tts generate \
    --text "Benchmarking pocket-tts performance." \
    --output benchmark.wav

The CLI prints metrics matching the format from line 701 of tts_model.py:


Generated: 1050 ms of audio in 250 ms so 4.20x faster than real-time

Python API and Metadata Retrieval

Access precise RTF and latency values programmatically via the return_metadata parameter:

from pocket_tts import TTSModel
from pocket_tts.default_parameters import DEFAULT_CONFIG

# Load model weights (~350 MiB RAM)

model = TTSModel.from_pretrained(DEFAULT_CONFIG)

# Generate with performance metadata

audio, meta = model.generate(
    "Testing CPU performance characteristics.", 
    return_metadata=True
)

print(f"RTF: {meta['rtf']:.2f}")           # e.g., 0.24

print(f"Latency: {meta['latency_ms']:.1f} ms")  # e.g., 240.0 ms

Real-Time Streaming Monitoring

Monitor per-chunk latency during streaming synthesis:

import sounddevice as sd
from pocket_tts import TTSModel

model = TTSModel.from_pretrained()
stream = model.generate_stream("Streaming example for latency testing")

for chunk in stream:
    # Each iteration logs timing to the default logger

    sd.play(chunk.numpy(), samplerate=model.sample_rate)
    sd.wait()

The streaming iterator invokes the same logging mechanism at line 701, allowing you to monitor RTF in real-time during interactive applications.

Summary

  • Real-time factor ranges from 0.17 to 0.45 on modern CPUs, translating to 2–6× faster-than-real-time synthesis.
  • Single-threaded execution via torch.set_num_threads(1) in tts_model.py ensures consistent latency but requires fast single-core performance.
  • Memory footprint remains under 400 MiB, with ~350 MiB for the model and minimal overhead for streaming buffers.
  • Latency per second of generated audio typically falls between 170 ms (Apple M4) and 450 ms (low-power laptop CPUs).
  • Performance metrics are accessible via CLI output, Python API metadata, or real-time logging during streaming generation.

Frequently Asked Questions

Does pocket-tts utilize multiple CPU cores for faster generation?

No. According to the source code in pocket_tts/models/tts_model.py, the implementation calls torch.set_num_threads(1) at line 49 by default, enforcing single-threaded execution. While experimental parallel decoding exists at line 557, the standard configuration relies entirely on single-core speed, meaning a 16-core CPU will not outperform a 4-core CPU with faster single-threaded performance.

What is the minimum RAM required to run pocket-tts?

Approximately 350 MiB of system RAM is required to load the Flow-LM and Mimi weights into memory. This makes pocket-tts feasible for most modern laptops, Raspberry Pi devices with sufficient RAM, and small cloud instances (t3.small or equivalent). Streaming generation adds only a few megabytes of temporary buffer space.

How does the real-time factor correlate with perceived latency?

The real-time factor (RTF) represents the ratio of audio duration to generation time. An RTF of 0.25 means 1 second of audio takes 250 ms to generate, creating a latency of 250 ms. Values below 1.0 indicate faster-than-real-time capability, which is essential for interactive applications. For responsive voice interfaces, target CPUs achieving RTF below 0.30 (approximately 300 ms latency per second of audio).

Can pocket-tts achieve real-time performance on ARM processors like Apple Silicon?

Yes. The MacBook Air M4 achieves an RTF of approximately 0.17 (6× real-time), demonstrating excellent performance on ARM architecture. Because the model runs on a single thread via PyTorch, it benefits from the high single-core performance and efficient memory architecture of Apple Silicon, making it well-suited for M-series Macs and compatible ARM servers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →