# How Voicebox Handles Paralinguistic Tags with Chatterbox Turbo: [laugh], [gasp], [sigh]

> Voicebox Chatterbox Turbo handles paralinguistic tags like [laugh] and [gasp] ensuring expressive sounds render correctly even in long content. Learn how.

- Repository: [Jamie Pine/voicebox](https://github.com/jamiepine/voicebox)
- Tags: deep-dive
- Published: 2026-04-14

---

**Voicebox passes raw paralinguistic tags such as `[laugh]`, `[gasp]`, and `[sigh]` directly to the Chatterbox Turbo TTS engine, while protecting tag integrity during chunked text processing to ensure expressive sounds render correctly even in long-form content.**

Voicebox is a modular speech-generation service that abstracts multiple TTS engines behind a common protocol. The repository `jamiepine/voicebox` implements the **Chatterbox Turbo** backend specifically to support these expressive markers, enabling the synthesis of emotional vocalizations like laughter and sighs alongside standard speech.

## Backend Architecture

All TTS engines in Voicebox implement the `TTSBackend` protocol defined in **[`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py)**. This protocol requires standardized methods including `load_model`, `generate`, and `unload_model` to ensure interoperability across different synthesis engines.

The system uses a factory pattern via `get_tts_backend_for_engine` to lazily instantiate concrete backend classes. When the requested engine equals `"chatterbox_turbo"`, the factory returns the specialized Chatterbox Turbo implementation and caches it for subsequent requests (see the `elif engine == "chatterbox_turbo"` branch around lines 62-66).

## Model Configuration

Chatterbox Turbo is configured through the `ModelConfig` dataclass, which specifies runtime requirements and constraints. The configuration declares English-only support (`languages=["en"]`) and flags the model as requiring post-generation audio trimming (`needs_trim=True`) to remove trailing artefacts common in the `ResembleAI/chatterbox-turbo` weights.

```python
ModelConfig(
    model_name="chatterbox-turbo",
    display_name="Chatterbox Turbo (English, Tags)",
    engine="chatterbox_turbo",
    hf_repo_id="ResembleAI/chatterbox-turbo",
    size_mb=1500,
    needs_trim=True,
    languages=["en"],
)

```

## Tag Processing in the Chatterbox Turbo Backend

The concrete implementation resides in **[`backend/backends/chatterbox_turbo_backend.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/chatterbox_turbo_backend.py)**. During model loading (`load_model`), the backend downloads weights from HuggingFace using `huggingface_hub.snapshot_download`, forces CPU execution on macOS, and applies a float-32 patch to ensure compatibility (lines 64-89).

The critical tag-handling logic occurs in the `generate` method (lines 160-162). This method forwards the raw input text—including any `[tag]` markers—directly to `self.model.generate` without preprocessing or stripping the brackets. The underlying `chatterbox-tts` library interprets these markers and synthesizes the corresponding paralinguistic sounds at the appropriate temporal positions. Tags function as inline instructions; they require no reference audio (`ref_audio`) to process, though the engine accepts voice prompts for speaker consistency when available.

## Chunked Generation with Tag Safety

For inputs exceeding optimal token lengths, Voicebox employs **chunked generation** via **[`backend/utils/chunked_tts.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/chunked_tts.py)** to prevent memory issues and quality degradation.

The text splitter (`split_text_into_chunks`) treats any `[...]` sequence as an indivisible atomic token using regex patterns (lines 56-71). The algorithm attempts sentence-level splits first, then clause boundaries, and finally whitespace. If forced to hard-split mid-text, the `_safe_hard_cut` function (lines 62-70) deliberately avoids cutting inside tag boundaries to prevent malformed markers like `[laug` or `h]`.

After parallel or sequential generation, `concatenate_audio_chunks` (lines 191-199) joins the resulting audio segments with a configurable crossfade (`crossfade_ms`) to eliminate click artifacts at chunk boundaries. This ensures that `[laugh]` or `[sigh]` tags positioned near split points remain intact and render audibly in the final output.

## Implementation Examples

### Single-Shot Generation with Tags

For short texts containing paralinguistic markers, call `generate` directly on the backend:

```python
import asyncio
from voicebox.backend.backends import get_tts_backend_for_engine

async def demo():
    backend = get_tts_backend_for_engine("chatterbox_turbo")
    await backend.load_model()
    
    voice_prompt = {"ref_audio": "path/to/reference.wav", "ref_text": "Hello world."}
    text = "Welcome to the show! [laugh] That was a great intro. [sigh]"
    
    audio, sample_rate = await backend.generate(text, voice_prompt)
    # audio is a NumPy float32 array; sample_rate is typically 24000 Hz

asyncio.run(demo())

```

### Long-Form Content with Tag Protection

For scripts exceeding `max_chunk_chars` (default 800), use the chunked utility to preserve tag integrity:

```python
import asyncio
from voicebox.backend.utils.chunked_tts import generate_chunked
from voicebox.backend.backends import get_tts_backend_for_engine

async def long_form_demo():
    backend = get_tts_backend_for_engine("chatterbox_turbo")
    await backend.load_model()
    
    long_script = ("This is a lengthy narration. " * 100) + "[gasp] I can't believe it!"
    
    audio, sr = await generate_chunked(
        backend,
        text=long_script,
        voice_prompt={},
        max_chunk_chars=800,
        crossfade_ms=50
    )
    # Tag is preserved and not split across chunks

asyncio.run(long_form_demo())

```

### API Endpoint Usage

Submit tagged text directly to the FastAPI endpoint:

```bash
curl -X POST http://localhost:8000/generate \
  -H "Content-Type: application/json" \
  -d '{
        "engine": "chatterbox_turbo",
        "text": "Oh really? [laugh] That is surprising! [cough]",
        "voice_prompt": null
      }' \
  --output expressive_speech.wav

```

## Summary

- **Voicebox** integrates Chatterbox Turbo through the `TTSBackend` protocol in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py).
- **Paralinguistic tags** (`[laugh]`, `[sigh]`, `[gasp]`, `[cough]`) pass raw to the model via the `generate` method in [`chatterbox_turbo_backend.py`](https://github.com/jamiepine/voicebox/blob/main/chatterbox_turbo_backend.py).
- **Chunked processing** in [`chunked_tts.py`](https://github.com/jamiepine/voicebox/blob/main/chunked_tts.py) treats tags as atomic units, preventing split markers during long-text generation.
- **English-only** support with mandatory post-generation trimming (`needs_trim=True`) ensures clean output.

## Frequently Asked Questions

### What paralinguistic tags does Chatterbox Turbo support?

The Chatterbox Turbo backend supports `[laugh]`, `[sigh]`, `[cough]`, and `[gasp]` markers. According to the docstring in [`chatterbox_turbo_backend.py`](https://github.com/jamiepine/voicebox/blob/main/chatterbox_turbo_backend.py) (lines 160-162), these tags instruct the model to inject non-verbal vocalizations at the specified temporal positions without requiring additional audio references.

### How does Voicebox prevent tags from being split during long text generation?

The `split_text_into_chunks` utility in [`backend/utils/chunked_tts.py`](https://github.com/jamiepine/voicebox/blob/main/backend/utils/chunked_tts.py) uses regex to identify `[...]` patterns as indivisible tokens. If forced to hard-split text, the `_safe_hard_cut` function (lines 62-70) avoids cutting inside tag boundaries, ensuring markers like `[laugh]` remain syntactically valid for the TTS model to interpret.

### Can I use Chatterbox Turbo with non-English text?

No. The `ModelConfig` in [`backend/backends/__init__.py`](https://github.com/jamiepine/voicebox/blob/main/backend/backends/__init__.py) explicitly sets `languages=["en"]` for Chatterbox Turbo. Attempting to synthesize non-English text will likely produce incorrect pronunciation or hallucinated tokens, as the underlying `ResembleAI/chatterbox-turbo` model was trained exclusively on English data.

### Does Chatterbox Turbo require reference audio to process tags?

No. Paralinguistic tags operate independently of voice cloning. While the `generate` method accepts an optional `voice_prompt` dictionary containing `ref_audio` for speaker consistency, tags like `[laugh]` process correctly without any reference file, as implemented in lines 160-162 of [`chatterbox_turbo_backend.py`](https://github.com/jamiepine/voicebox/blob/main/chatterbox_turbo_backend.py).