# How SentencePiece Tokenization Works in Pocket-TTS LUTConditioner

> Understand SentencePiece tokenization in Pocket-TTS LUTConditioner. Learn how it maps text to integer IDs for embedding vectors with n_bins + 1 lookup table entries.

- Repository: [kyutai/pocket-tts](https://github.com/kyutai-labs/pocket-tts)
- Tags: deep-dive
- Published: 2026-07-09

---

**The `LUTConditioner` module converts raw text into embedding vectors using a SentencePiece sub-word tokenizer that maps strings to integer IDs via a lookup table with `n_bins + 1` entries.**

**LUTConditioner** serves as the text-conditioning bridge in Pocket-TTS, transforming input strings into dense embeddings that drive the flow-based language model. According to the Kyutai Labs source code, this conversion relies on **SentencePiece**, a language-agnostic sub-word tokenization algorithm that ships with the repository. The implementation wraps the tokenizer in a dedicated class and pairs it with an `nn.Embedding` layer to produce sequence representations.

## Tokenization Pipeline Architecture

The `LUTConditioner` follows a strict two-phase design that separates tokenization from embedding lookup. This architecture allows the model to cache pre-tokenized inputs while maintaining a clean forward-pass interface.

### Tokenizer Initialization

When instantiating `LUTConditioner`, the module receives `n_bins` (defining the vocabulary size) and a path to a SentencePiece model file. In [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py), the constructor initializes the **SentencePieceTokenizer** wrapper:

```python
self.tokenizer = SentencePieceTokenizer(n_bins, tokenizer_path)

```

The wrapper loads the underlying `sentencepiece.SentencePieceProcessor` and handles model downloading via `download_if_necessary` if the file is not present locally. By default, Pocket-TTS uses a vocabulary size of 4000 tokens (`DEFAULT_TOKENIZER_N_BINS`), ensuring a balance between granularity and embedding table efficiency.

### Text Encoding and Embedding Lookup

The `prepare()` method orchestrates the conversion from string to tensor. When called, it invokes the tokenizer’s `__call__` method, which executes `sp.encode(text, out_type=int)` and wraps the result in a `torch.Tensor` with shape `[1, seq_len]`:

```python
tokens = self.tokenizer(x)  # Returns TokenizedText tuple

tokens = tokens[0].to(self.embed.weight.device)

```

These integer IDs are then fed into an `nn.Embedding` layer defined with dimensions `(n_bins + 1, dim)`. The `+1` reserves index 0 for padding, ensuring that every valid token ID (1-indexed) maps to a learnable embedding vector. The output tensor of shape `[seq_len, dim]` represents the final conditioner output that drives the TTS model.

## Implementation Details

### SentencePieceTokenizer Wrapper

Located in [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py), the `SentencePieceTokenizer` class abstracts the underlying SentencePiece processor. It manages device placement and batch dimension handling, ensuring compatibility with PyTorch training loops. The tokenizer downloads the model file from HuggingFace on first use if the path is not already cached locally.

### Base Conditioner Contract

`LUTConditioner` inherits from `BaseConditioner` (defined in [`pocket_tts/conditioners/base.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/base.py)), which standardizes the conditioner interface across the codebase. The base class `forward()` simply delegates to `_get_condition`, allowing subclasses to implement custom embedding logic while maintaining consistent API signatures. This design enables the `TTSModel` to treat text conditioners interchangeably with other conditioning modalities.

### Embedding Layer Configuration

The embedding table dimensions are critical for model stability. With `n_bins` set to 4000 and `dim` typically matching the model’s hidden size, the layer contains over one million parameters. The padding index at 0 remains untrained during inference, serving as a buffer for batch collation.

## Practical Usage Examples

### Basic Text Conditioning

```python
from pocket_tts.conditioners.text import LUTConditioner, get_default_tokenizer

# Initialize the default SentencePiece tokenizer (4000 vocab)

tokenizer = get_default_tokenizer()

# Create conditioner with 4000 bins, 256-dimensional embeddings

conditioner = LUTConditioner(
    n_bins=4000,
    tokenizer_path=tokenizer.sp.model_file(),
    dim=256,
    output_dim=256,
)

# Convert text to embeddings

text = "Hello, Pocket-TTS!"
tokenized = conditioner.prepare(text)
embeddings = conditioner(tokenized)
print(embeddings.shape)  # Output: torch.Size([seq_len, 256])

```

### Integration with TTSModel

```python
from pocket_tts.models.tts_model import TTSModel

# Instantiate the complete TTS pipeline

tts = TTSModel()

# The conditioner can be passed directly to the generation method

audio_stream = tts.generate_audio_stream(
    text_conditioner=conditioner,
    text="The quick brown fox jumps over the lazy dog.",
)

```

## Summary

- **SentencePiece integration** in `LUTConditioner` uses a 4000-token sub-word vocabulary to handle multilingual text input efficiently.
- **Two-phase processing** separates `prepare()` (tokenization) from `forward()` (embedding lookup), enabling flexible inference patterns.
- **Embedding dimensions** are set to `(n_bins + 1, dim)` to accommodate padding index 0, with the default configuration using 4001 rows.
- **Source files** implementing this logic reside in [`pocket_tts/conditioners/text.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/text.py) and [`pocket_tts/conditioners/base.py`](https://github.com/kyutai-labs/pocket-tts/blob/main/pocket_tts/conditioners/base.py).

## Frequently Asked Questions

### What vocabulary size does Pocket-TTS use for SentencePiece tokenization?

The default configuration uses **4000 tokens** (`DEFAULT_TOKENIZER_N_BINS`), which provides sufficient granularity for multilingual speech synthesis while keeping the embedding table size manageable. This value is passed as `n_bins` to the `LUTConditioner` constructor.

### How does LUTConditioner handle out-of-vocabulary words?

SentencePiece decomposes rare words into sub-word units, effectively eliminating out-of-vocabulary issues. The 4000-token vocabulary covers common character sequences and morphemes, allowing the tokenizer to represent any input string as a sequence of known sub-word IDs.

### Can I use a custom SentencePiece model with LUTConditioner?

Yes. While `get_default_tokenizer()` provides the pretrained model, you can instantiate `LUTConditioner` with any valid SentencePiece model file by specifying the `tokenizer_path` parameter. Ensure your custom model’s vocabulary size matches the `n_bins` parameter to avoid indexing errors in the embedding layer.

### Why does the embedding table have n_bins + 1 entries instead of n_bins?

The additional entry at index 0 serves as a **padding token**. This allows the model to process variable-length sequences in batches, where shorter sequences are padded to the maximum length in the batch. The `+1` ensures that valid token IDs (which start at 1) do not collide with the padding index.