# How Pocket-TTS Handles Text Tokenization with SentencePiece and LUTConditioner

> Learn how Pocket-TTS tokenizes text using SentencePiece and LUTConditioner. Discover how raw text becomes dense embedding sequences for efficient model processing.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: deep-dive
- Published: 2026-07-11

---

**Pocket-TTS converts raw text into dense embedding sequences using a `SentencePieceTokenizer` to generate integer token IDs and a `LUTConditioner` to map those IDs to learned vectors via an `nn.Embedding` lookup table.**

Pocket-TTS, the open-source text-to-speech engine from Kyutai Labs, isolates linguistic preprocessing in its conditioners package. The system implements **text tokenization with SentencePiece and LUTConditioner** through two tightly-coupled classes that bridge natural language input and the neural audio generator. This architecture lives in [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py) and leverages SentencePiece subword segmentation combined with learnable lookup-table embeddings.

## The SentencePieceTokenizer Implementation

### Subword Tokenization and Vocab Management

The `SentencePieceTokenizer` class (defined at lines 13–36 in [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py)) wraps a pre-trained SentencePiece model to handle string-to-integer conversion. During initialization, it downloads the tokenizer asset via `download_if_necessary` (line 27) and instantiates a `sentencepiece.SentencePieceProcessor` (line 29).

A critical validation occurs at lines 30–32 where the code asserts that the supplied `n_bins` parameter exactly matches the SentencePiece vocabulary size. This ensures the downstream embedding layer allocates the correct number of rows.

When called, the tokenizer executes:

```python
return TokenizedText(torch.tensor(self.sp.encode(text, out_type=int))[None, :])

```

This operation (line 35) returns a `TokenizedText` container wrapping a 2-D tensor of shape `[1, sequence_length]` containing the token IDs.

## The LUTConditioner Embedding Layer

### Architecture and Initialization

`LUTConditioner` (lines 53–77 in [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py)) inherits from `BaseConditioner` and manages the embedding lookup. Its constructor creates two essential components:

- A `SentencePieceTokenizer` instance (line 66)
- An `nn.Embedding` layer with `n_bins + 1` rows and `dim` columns (line 67)

The extra row (+1) accommodates padding indices, while `dim` specifies the vector size (commonly 256 in the default configuration).

### From Text to Tensors

The `prepare` method (lines 69–72) handles device placement and batching:

1. Tokenizes the input string using the embedded tokenizer
2. Moves the resulting IDs to the same device as the embedding weights
3. Returns a `TokenizedText` object ready for conditioning

The `_get_condition` method (lines 75–76) performs the actual embedding lookup:

```python
return self.embed(text.tokens)

```

Here, `text.tokens` contains the integer IDs, and the method returns a tensor of shape `[sequence_length, dim]` representing the dense text condition fed into the TTS model.

## End-to-End Tokenization Pipeline

The complete **text tokenization with SentencePiece and LUTConditioner** flow follows three distinct stages:

1. **Model Loading**: `SentencePieceTokenizer` loads the pre-trained subword model (default path: `hf://kyutai/pocket-tts-without-voice-cloning/tokenizer.model`)
2. **ID Generation**: Raw strings convert to integer sequences via `sp.encode()`
3. **Embedding Projection**: `LUTConditioner` projects each ID to a dense vector using the trainable `nn.Embedding` table

This separation allows the core TTS model in [`pocket_tts/models/tts_model.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/models/tts_model.py) to consume fixed-dimensional conditions without handling raw string parsing.

## Practical Implementation Examples

### Standard Conditioning Workflow

To generate embeddings for a TTS forward pass:

```python
from pocket_tts.conditioners.text import get_default_tokenizer, LUTConditioner

# Initialize default tokenizer (4000 vocab entries)

tokenizer = get_default_tokenizer()

# Create conditioner with 256-dimensional embeddings

conditioner = LUTConditioner(
    n_bins=4000,
    tokenizer_path=tokenizer.sp.model(),
    dim=256,
    output_dim=256,
)

# Prepare text (handles tokenization and device placement)

tokenized = conditioner.prepare("Hello, Pocket-TTS!")

# Extract condition tensor

condition_tensor = conditioner._get_condition(tokenized)
print(condition_tensor.shape)  # torch.Size([6, 256])

```

### Direct Token ID Access

For debugging or custom preprocessing, access the raw SentencePiece tokenizer directly:

```python
ids = tokenizer("Direct tokenization example")
print(ids[0])  # tensor([ 12, 345, 78, ...])

```

## Summary

- **`SentencePieceTokenizer`** in [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py) handles subword segmentation, validating that `n_bins` matches the SentencePiece vocabulary size.
- **`LUTConditioner`** manages an `nn.Embedding` table with `n_bins + 1` rows, converting token IDs to dense vectors via `prepare()` and `_get_condition()`.
- The pipeline downloads tokenizer assets automatically using `download_if_necessary` from [`pocket_tts/utils/utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/utils.py).
- Default configuration uses a 4000-token vocabulary with 256-dimensional embeddings.

## Frequently Asked Questions

### What is the relationship between SentencePieceTokenizer and LUTConditioner?

`SentencePieceTokenizer` performs the initial string-to-integer conversion using subword segmentation, while `LUTConditioner` contains the tokenizer instance and an embedding layer that converts those integers into trainable vectors. The conditioner calls the tokenizer internally during its `prepare` method, making `LUTConditioner` the higher-level interface for the TTS model.

### Why does the embedding table have n_bins + 1 rows instead of n_bins?

The extra row provides a dedicated padding index (typically index 0) for batched sequences of variable length. This allows the model to distinguish between valid vocabulary tokens (indices 1 to n_bins) and padding positions during attention and convolution operations.

### Where does Pocket-TTS download the SentencePiece model from?

By default, the tokenizer downloads from `hf://kyutai/pocket-tts-without-voice-cloning/tokenizer.model` via the `download_if_necessary` utility in [`pocket_tts/utils/utils.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/utils/utils.py). This occurs during `SentencePieceTokenizer` initialization at line 27 of [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py).