# Pocket-tts Performance Characteristics on CPUs: RTF, Latency, and Memory Usage

> Discover Pocket-tts performance on CPUs. Explore RTF, latency, and memory usage benchmarks. Achieve 2-6x faster audio generation with low RAM consumption.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: performance
- Published: 2026-07-11

---

**Pocket-tts achieves real-time factors between 0.17 and 0.45 on modern CPUs, generating audio 2–6× faster than real-time while consuming approximately 350 MiB of RAM.**

Pocket-tts from kyutai-labs is an open-source text-to-speech system optimized for **CPU-only deployment**. The architecture combines a **Flow-LM** generative model with a **Mimi** neural audio codec to deliver streaming speech synthesis without GPU acceleration. Because the entire pipeline runs in single-threaded mode by default, performance scales directly with single-core processor speed, making it ideal for edge devices and modest cloud instances.

## Real-Time Factor and Latency Architecture

### Single-Threaded Execution Model

In [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py), the model explicitly configures **single-threaded PyTorch execution** to ensure deterministic performance. At line 49, the code calls:

```python
torch.set_num_threads(1)

```

This design choice means pocket-tts will not automatically utilize multiple CPU cores during generation. While line 557 contains experimental parallel decoding logic, the default behavior relies entirely on **single-core speed**. Consequently, generation performance scales with your CPU's single-threaded performance rather than core count.

### RTF Calculation and Logging

The repository automatically calculates the **real-time factor (RTF)** during generation. In [`tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/tts_model.py) at line 701, the code logs performance metrics for each generated chunk:

```python
logger.info(
    "Generated: %d ms of audio in %d ms so %.2fx faster than real-time",
    audio_ms, generation_ms, audio_ms / generation_ms
)

```

This produces output like `6.0x faster than real-time`, indicating an RTF of approximately 0.17.

### CPU Benchmark Comparison

Because pocket-tts is single-threaded, performance varies by processor single-core speed. Community benchmarks and author notes indicate the following characteristics:

| CPU (Single-Core) | Approximate RTF | Latency per 1s Audio | RAM Usage |
|-------------------|-----------------|---------------------|-----------|
| **Apple M4** (MacBook Air 2023) | 0.17 (6× RTF) | ~170 ms | ~350 MiB |
| **Intel i7-12700H** (~3.5 GHz) | 0.22 (4.5× RTF) | ~220 ms | ~350 MiB |
| **AMD Ryzen 7 5800X** (~3.8 GHz) | 0.20 (5× RTF) | ~200 ms | ~350 MiB |
| **Intel i5-8250U** (~1.6 GHz, low-end laptop) | 0.45 (2.2× RTF) | ~450 ms | ~350 MiB |

*RTF represents audio duration divided by generation time; lower values indicate faster-than-real-time performance.*

## Memory Usage Characteristics

### Model Weights and RAM Footprint

The **Flow-LM** and **Mimi** components together occupy approximately **300 MiB** for model weights. When loaded in evaluation mode via `model.eval()`, the complete memory footprint reaches roughly **350 MiB** on x86-64 and ARM64 systems. This modest requirement allows deployment on laptops with limited RAM and small cloud instances.

### Streaming Buffer Overhead

During streaming generation, intermediate tensors and audio buffers add minimal overhead. Each frame represents approximately **0.08 seconds** of audio, requiring only **~10 kB** of additional memory. The streaming state therefore adds only a few megabytes to the base footprint, enabling continuous synthesis on memory-constrained devices.

## Measuring Performance Programmatically

### Command-Line Interface Metrics

Install pocket-tts and generate speech to see real-time performance statistics printed to stderr:

```bash

# Install with quantization support

uv pip install "pocket-tts[quantize]"

# Generate audio and observe RTF metrics

uv run pocket-tts generate \
    --text "Benchmarking pocket-tts performance." \
    --output benchmark.wav

```

The CLI prints metrics matching the format from line 701 of [`tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/tts_model.py):

```

Generated: 1050 ms of audio in 250 ms so 4.20x faster than real-time

```

### Python API and Metadata Retrieval

Access precise RTF and latency values programmatically via the `return_metadata` parameter:

```python
from pocket_tts import TTSModel
from pocket_tts.default_parameters import DEFAULT_CONFIG

# Load model weights (~350 MiB RAM)

model = TTSModel.from_pretrained(DEFAULT_CONFIG)

# Generate with performance metadata

audio, meta = model.generate(
    "Testing CPU performance characteristics.", 
    return_metadata=True
)

print(f"RTF: {meta['rtf']:.2f}")           # e.g., 0.24

print(f"Latency: {meta['latency_ms']:.1f} ms")  # e.g., 240.0 ms

```

### Real-Time Streaming Monitoring

Monitor per-chunk latency during streaming synthesis:

```python
import sounddevice as sd
from pocket_tts import TTSModel

model = TTSModel.from_pretrained()
stream = model.generate_stream("Streaming example for latency testing")

for chunk in stream:
    # Each iteration logs timing to the default logger

    sd.play(chunk.numpy(), samplerate=model.sample_rate)
    sd.wait()

```

The streaming iterator invokes the same logging mechanism at line 701, allowing you to monitor RTF in real-time during interactive applications.

## Summary

- **Real-time factor** ranges from **0.17 to 0.45** on modern CPUs, translating to 2–6× faster-than-real-time synthesis.
- **Single-threaded execution** via `torch.set_num_threads(1)` in [`tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/tts_model.py) ensures consistent latency but requires fast single-core performance.
- **Memory footprint** remains under **400 MiB**, with ~350 MiB for the model and minimal overhead for streaming buffers.
- **Latency** per second of generated audio typically falls between **170 ms** (Apple M4) and **450 ms** (low-power laptop CPUs).
- Performance metrics are accessible via CLI output, Python API metadata, or real-time logging during streaming generation.

## Frequently Asked Questions

### Does pocket-tts utilize multiple CPU cores for faster generation?

No. According to the source code in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py), the implementation calls `torch.set_num_threads(1)` at line 49 by default, enforcing single-threaded execution. While experimental parallel decoding exists at line 557, the standard configuration relies entirely on single-core speed, meaning a 16-core CPU will not outperform a 4-core CPU with faster single-threaded performance.

### What is the minimum RAM required to run pocket-tts?

Approximately **350 MiB** of system RAM is required to load the Flow-LM and Mimi weights into memory. This makes pocket-tts feasible for most modern laptops, Raspberry Pi devices with sufficient RAM, and small cloud instances (t3.small or equivalent). Streaming generation adds only a few megabytes of temporary buffer space.

### How does the real-time factor correlate with perceived latency?

The **real-time factor (RTF)** represents the ratio of audio duration to generation time. An RTF of 0.25 means 1 second of audio takes 250 ms to generate, creating a **latency** of 250 ms. Values below 1.0 indicate faster-than-real-time capability, which is essential for interactive applications. For responsive voice interfaces, target CPUs achieving RTF below 0.30 (approximately 300 ms latency per second of audio).

### Can pocket-tts achieve real-time performance on ARM processors like Apple Silicon?

Yes. The **MacBook Air M4** achieves an RTF of approximately **0.17** (6× real-time), demonstrating excellent performance on ARM architecture. Because the model runs on a single thread via PyTorch, it benefits from the high single-core performance and efficient memory architecture of Apple Silicon, making it well-suited for M-series Macs and compatible ARM servers.