How to Use Inline Emotion and Prosody Control Tags in Fish-Speech
Insert natural-language descriptions inside square brackets—such as [laugh] or [whispers]—directly into your text to change speaking style during synthesis.
Fish-Speech (fishaudio/fish-speech) supports inline emotion and prosody control tags that let you manipulate vocal expression without training separate models. By wrapping descriptors in square brackets, you can trigger acoustic effects at specific positions within a single utterance.
Understanding Inline Control Tags
Fish-Speech interprets any word or phrase wrapped in square brackets ([ … ]) as a prosody instruction. During tokenization, the brackets and their contents are preserved as a single token, allowing the model to map that token to a distinct acoustic embedding that drives the decoder.
The model recognizes both standard tags and free-form descriptions:
[laugh]– Adds a short laugh after the preceding word[whispers]– Switches to a low-volume, breathy voice[super happy]– Produces brighter, higher-pitched, energetic speech[whisper in small voice]or[professional broadcast tone]– Custom descriptions work if semantically similar to training data
Placement Rules for Maximum Effect
Inline Positioning
Place the tag exactly where you want the effect to start. The model applies the change immediately after the token preceding the tag.
Hello [laugh] world!
Chaining Multiple Tags
You can insert several tags in sequence. When tags overlap, the later tag overrides the previous effect.
This is amazing [super happy] but now I'm [whispers] telling a secret.
Boundary Placement
A tag at the beginning of the utterance affects the entire sentence. A tag at the end has no acoustic effect because no speech follows it.
How the Tokenizer Preserves Tags
The preservation of bracket tags happens in fish_speech/tokenizer.py. The FishTokenizer.encode() method forces the underlying HuggingFace tokenizer to keep all special tokens by setting allowed_special="all" (lines 8–14). Lines 8–16 define the special tokens vocabulary, ensuring that square-bracket strings remain intact as individual tokens rather than being split into subwords.
from fish_speech.tokenizer import FishTokenizer
tokenizer = FishTokenizer("fishaudio/s2-pro")
input_ids = tokenizer.encode("Hello [laugh] world", add_special_tokens=False)
# Tags are preserved as separate tokens in the input_ids list
Inference Pipeline Integration
After tokenization, the list—including tag tokens—is forwarded to the semantic model. In fish_speech/inference_engine/__init__.py, the inference method passes these tokens to the LLaMA-style text-to-semantic model. The model learns a direct mapping from each tag token to a prosody embedding, which subsequently conditions the acoustic decoder to render the requested emotion or speaking style.
Practical Implementation Examples
Python API Usage
Use the FishSpeech class alongside FishTokenizer to process tagged text programmatically:
from fish_speech import FishSpeech
from fish_speech.tokenizer import FishTokenizer
# Load model and tokenizer
tts = FishSpeech.from_pretrained("fishaudio/s2-pro")
tokenizer = FishTokenizer("fishaudio/s2-pro")
# Text with inline tags
text = "Hey there [laugh] how are you doing today [whispers]?"
# Encode preserves tags as tokens
input_ids = tokenizer.encode(text, add_special_tokens=False)
# Generate audio
audio = tts.infer(input_ids)
tts.save_wav(audio, "output.wav")
Command-Line Interface
The bundled CLI forwards raw strings directly to the tokenizer, making tags work without additional configuration:
fish-speech \
--model checkpoints/s2-pro \
--text "Good morning [super happy] everyone! [whispers] This is a secret." \
--output output.wav
Best Practices and Limitations
- Exact spelling required – The token must match the training vocabulary exactly (e.g.,
[laugh]works;[laughs]does not). - No nested brackets – Avoid
[laugh [whispers]]; the tokenizer will split or ignore invalid nested structures. - Experiment with custom tags – Descriptions like
[pitch up]may produce plausible effects if semantically close to known tags. - Speaker changes use different syntax – Multi-speaker control uses
<|speaker:i|>tokens, not square brackets.
Summary
- Wrap emotion or prosody cues in square brackets (e.g.,
[whispers],[super happy]) and insert them inline where the effect should begin. - The tokenizer in
fish_speech/tokenizer.pypreserves these brackets as single tokens by settingallowed_special="all". - The inference engine in
fish_speech/inference_engine/__init__.pymaps tag tokens to acoustic embeddings that drive the decoder. - Multiple tags can chain together, with later tags overriding earlier ones.
- Tags must match training spellings exactly; custom descriptions work when semantically similar to learned tags.
Frequently Asked Questions
What happens if I misspell an emotion tag?
The model treats misspelled tags as unknown tokens. Because the tokenizer in fish_speech/tokenizer.py requires exact vocabulary matches when allowed_special="all" is set, a tag like [laughs] will likely be split into subword tokens and ignored as a prosody cue, resulting in standard speech without the intended effect.
Can I use multiple emotion tags in the same sentence?
Yes. You can chain multiple tags within a single utterance. When tags overlap, the later tag overrides the previous prosody setting. For example, [super happy] [whispers] will end with a whispered tone, not a happy one.
Do custom prosody descriptions work with any text?
Custom descriptions work if the wording is semantically similar to tags seen during training. While standard tags like [laugh] are guaranteed to work, free-form descriptions such as [professional broadcast tone] rely on the model's ability to generalize from learned prosody embeddings. Results may vary based on how close the description is to the training distribution.
Why doesn't a tag at the end of the text produce any effect?
The model applies prosody changes immediately after the token preceding the tag. If a tag appears at the end of the text with no following speech, there is no audio segment to apply the effect to. Place tags before the text you want affected, or at the very beginning to influence the entire utterance.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →