How Voicebox Handles Paralinguistic Tags with Chatterbox Turbo: [laugh], [gasp], [sigh]
Voicebox passes raw paralinguistic tags such as [laugh], [gasp], and [sigh] directly to the Chatterbox Turbo TTS engine, while protecting tag integrity during chunked text processing to ensure expressive sounds render correctly even in long-form content.
Voicebox is a modular speech-generation service that abstracts multiple TTS engines behind a common protocol. The repository jamiepine/voicebox implements the Chatterbox Turbo backend specifically to support these expressive markers, enabling the synthesis of emotional vocalizations like laughter and sighs alongside standard speech.
Backend Architecture
All TTS engines in Voicebox implement the TTSBackend protocol defined in backend/backends/__init__.py. This protocol requires standardized methods including load_model, generate, and unload_model to ensure interoperability across different synthesis engines.
The system uses a factory pattern via get_tts_backend_for_engine to lazily instantiate concrete backend classes. When the requested engine equals "chatterbox_turbo", the factory returns the specialized Chatterbox Turbo implementation and caches it for subsequent requests (see the elif engine == "chatterbox_turbo" branch around lines 62-66).
Model Configuration
Chatterbox Turbo is configured through the ModelConfig dataclass, which specifies runtime requirements and constraints. The configuration declares English-only support (languages=["en"]) and flags the model as requiring post-generation audio trimming (needs_trim=True) to remove trailing artefacts common in the ResembleAI/chatterbox-turbo weights.
ModelConfig(
model_name="chatterbox-turbo",
display_name="Chatterbox Turbo (English, Tags)",
engine="chatterbox_turbo",
hf_repo_id="ResembleAI/chatterbox-turbo",
size_mb=1500,
needs_trim=True,
languages=["en"],
)
Tag Processing in the Chatterbox Turbo Backend
The concrete implementation resides in backend/backends/chatterbox_turbo_backend.py. During model loading (load_model), the backend downloads weights from HuggingFace using huggingface_hub.snapshot_download, forces CPU execution on macOS, and applies a float-32 patch to ensure compatibility (lines 64-89).
The critical tag-handling logic occurs in the generate method (lines 160-162). This method forwards the raw input text—including any [tag] markers—directly to self.model.generate without preprocessing or stripping the brackets. The underlying chatterbox-tts library interprets these markers and synthesizes the corresponding paralinguistic sounds at the appropriate temporal positions. Tags function as inline instructions; they require no reference audio (ref_audio) to process, though the engine accepts voice prompts for speaker consistency when available.
Chunked Generation with Tag Safety
For inputs exceeding optimal token lengths, Voicebox employs chunked generation via backend/utils/chunked_tts.py to prevent memory issues and quality degradation.
The text splitter (split_text_into_chunks) treats any [...] sequence as an indivisible atomic token using regex patterns (lines 56-71). The algorithm attempts sentence-level splits first, then clause boundaries, and finally whitespace. If forced to hard-split mid-text, the _safe_hard_cut function (lines 62-70) deliberately avoids cutting inside tag boundaries to prevent malformed markers like [laug or h].
After parallel or sequential generation, concatenate_audio_chunks (lines 191-199) joins the resulting audio segments with a configurable crossfade (crossfade_ms) to eliminate click artifacts at chunk boundaries. This ensures that [laugh] or [sigh] tags positioned near split points remain intact and render audibly in the final output.
Implementation Examples
Single-Shot Generation with Tags
For short texts containing paralinguistic markers, call generate directly on the backend:
import asyncio
from voicebox.backend.backends import get_tts_backend_for_engine
async def demo():
backend = get_tts_backend_for_engine("chatterbox_turbo")
await backend.load_model()
voice_prompt = {"ref_audio": "path/to/reference.wav", "ref_text": "Hello world."}
text = "Welcome to the show! [laugh] That was a great intro. [sigh]"
audio, sample_rate = await backend.generate(text, voice_prompt)
# audio is a NumPy float32 array; sample_rate is typically 24000 Hz
asyncio.run(demo())
Long-Form Content with Tag Protection
For scripts exceeding max_chunk_chars (default 800), use the chunked utility to preserve tag integrity:
import asyncio
from voicebox.backend.utils.chunked_tts import generate_chunked
from voicebox.backend.backends import get_tts_backend_for_engine
async def long_form_demo():
backend = get_tts_backend_for_engine("chatterbox_turbo")
await backend.load_model()
long_script = ("This is a lengthy narration. " * 100) + "[gasp] I can't believe it!"
audio, sr = await generate_chunked(
backend,
text=long_script,
voice_prompt={},
max_chunk_chars=800,
crossfade_ms=50
)
# Tag is preserved and not split across chunks
asyncio.run(long_form_demo())
API Endpoint Usage
Submit tagged text directly to the FastAPI endpoint:
curl -X POST http://localhost:8000/generate \
-H "Content-Type: application/json" \
-d '{
"engine": "chatterbox_turbo",
"text": "Oh really? [laugh] That is surprising! [cough]",
"voice_prompt": null
}' \
--output expressive_speech.wav
Summary
- Voicebox integrates Chatterbox Turbo through the
TTSBackendprotocol inbackend/backends/__init__.py. - Paralinguistic tags (
[laugh],[sigh],[gasp],[cough]) pass raw to the model via thegeneratemethod inchatterbox_turbo_backend.py. - Chunked processing in
chunked_tts.pytreats tags as atomic units, preventing split markers during long-text generation. - English-only support with mandatory post-generation trimming (
needs_trim=True) ensures clean output.
Frequently Asked Questions
What paralinguistic tags does Chatterbox Turbo support?
The Chatterbox Turbo backend supports [laugh], [sigh], [cough], and [gasp] markers. According to the docstring in chatterbox_turbo_backend.py (lines 160-162), these tags instruct the model to inject non-verbal vocalizations at the specified temporal positions without requiring additional audio references.
How does Voicebox prevent tags from being split during long text generation?
The split_text_into_chunks utility in backend/utils/chunked_tts.py uses regex to identify [...] patterns as indivisible tokens. If forced to hard-split text, the _safe_hard_cut function (lines 62-70) avoids cutting inside tag boundaries, ensuring markers like [laugh] remain syntactically valid for the TTS model to interpret.
Can I use Chatterbox Turbo with non-English text?
No. The ModelConfig in backend/backends/__init__.py explicitly sets languages=["en"] for Chatterbox Turbo. Attempting to synthesize non-English text will likely produce incorrect pronunciation or hallucinated tokens, as the underlying ResembleAI/chatterbox-turbo model was trained exclusively on English data.
Does Chatterbox Turbo require reference audio to process tags?
No. Paralinguistic tags operate independently of voice cloning. While the generate method accepts an optional voice_prompt dictionary containing ref_audio for speaker consistency, tags like [laugh] process correctly without any reference file, as implemented in lines 160-162 of chatterbox_turbo_backend.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →