How to Implement Controllable TTS for Agent Voice Output Using Fish Audio S1

Controllable TTS enables dynamic manipulation of emotional tone, speaking speed, and conversational style by selecting from a curated library of 24 reference voice clips and parsing inline markup tags.

The bojieli/ai-agent-book repository provides a complete implementation of controllable text-to-speech built on Fish Audio S1. This architecture allows AI agents to express happiness, frustration, or thinking states while adjusting tempo and formality on demand. The system decomposes the synthesis pipeline into three specialized layers that work together to transform marked-up scripts into expressive audio.

The Three-Layer Architecture

The implementation separates concerns into distinct modules that handle voice storage, script interpretation, and audio generation.

Reference Library

The voice library stores 24 real-voice clips covering every combination of emotion, speed, and style. Located in chapter6/controllable-tts/voice_library.py, this module manages the manifest that maps prosody profiles to reference audio files. The library handles neutral, happy, frustrated, and thinking emotions crossed with slow, normal, and fast speeds, further divided into formal and casual styles.

Markup Parser

The parser in chapter6/controllable-tts/markup.py converts annotated scripts into structured segments. It processes state tags like [EMO:happy], [SPEED:fast], and [STYLE:casual] that persist until changed, and inline tags such as [THINKING], <pause>, or [BREATH] that insert non-verbal cues or silences.

Synthesis Engine

The tts.py module orchestrates voice cloning via the Fish Audio SDK. For each speech segment, it selects the appropriate reference clip based on current markup state, calls synth_speech to generate audio, and concatenates results using FFmpeg. Silence segments are synthesized separately and merged into the final MP3.

Setting Up the Reference Library

Before parsing scripts, you must build the 24-clip library that enables prosody control. This requires a Fish API key and a base reference ID representing your owned voice.

Run the build script from the repository root:

export FISH_API_KEY="your_api_key_here"
python -m chapter6.controllable-tts.build_reference_library \
    --base-reference-id <YOUR_BASE_REFERENCE_ID>

This creates a reference_audio/ directory containing the clips and a manifest.json file that voice_library.py uses to resolve prosody requests.

Writing and Parsing Control Markup

Scripts use a lightweight markup syntax to specify voice characteristics. State tags modify all subsequent text until overridden, while inline tags create immediate effects.

from chapter6.controllable_tts.markup import parse

demo_script = (
    "[EMO:happy][SPEED:fast][STYLE:casual]太好了!您的订单已确认。"
    "[THINKING]让我查一下发货时间。"
    "[EMO:neutral][SPEED:normal][STYLE:formal]预计明天下午送达。"
)

segments = parse(demo_script)
print(segments)   # List of Segment objects (speech or silence)

Supported tags include:

  • Emotion tags: [EMO:neutral], [EMO:happy], [EMO:frustrated], [EMO:thinking]
  • Speed tags: [SPEED:slow], [SPEED:normal], [SPEED:fast]
  • Style tags: [STYLE:formal], [STYLE:casual]
  • Non-verbal markers: [THINKING], [BREATH], <pause>

Synthesizing Speech with Prosody Control

With segments parsed and the library built, the synthesize_segments function in chapter6/controllable-tts/tts.py handles the full audio pipeline.

from chapter6.controllable_tts.tts import synthesize_segments
from chapter6.controllable_tts.voice_library import load_voice_library, DEFAULT_MANIFEST

# Load the manifest created during the build step

library = load_voice_library(DEFAULT_MANIFEST)

# Generate controlled audio

output_path = "output/agent_output.mp3"
info = synthesize_segments(
    segments=segments,
    out_path=output_path,
    workdir="tmp_workdir",
    manifest_path=DEFAULT_MANIFEST,
)

print(f"Audio written to {output_path}")

For scenarios requiring only basic voice cloning without prosody shifting, use synth_direct_reference:

from chapter6.controllable_tts.tts import synth_direct_reference

plain_text = "Hello, this is a plain TTS request."
synth_direct_reference(
    text=plain_text,
    reference_id=library["source_reference_id"],
    out_path="output/plain_tts.mp3",
)

Running the Complete Demo

The repository includes demo.py to compare synthesis methods side-by-side. It generates three variants of the same input to demonstrate the value of controllable TTS.

python -m chapter6.controllable_tts.demo \
    --text "[EMO:happy][SPEED:fast][STYLE:casual]你好!" \
    --output-dir ./demo_output

This produces:

  • A_no_control_markers.mp3: Direct S1 synthesis without reference library control
  • B_single_reference.mp3: Single-clip reference (neutral + normal + formal)
  • C_24_reference_library.mp3: Fully controllable output using the complete 24-clip library

The command also writes an evidence.json file documenting the experimental configuration and segment metadata.

Summary

  • Controllable TTS requires a multi-clip reference library covering emotion, speed, and style dimensions.
  • The voice_library.py module manages 24 distinct reference clips indexed by prosody profile.
  • markup.py parses control tags into segments that drive reference selection during synthesis.
  • tts.py handles Fish Audio SDK integration, audio generation, and FFmpeg-based concatenation.
  • The system supports both fully controlled synthesis via synthesize_segments and basic cloning via synth_direct_reference.

Frequently Asked Questions

What is the minimum number of reference clips needed for basic controllable TTS?

The implementation uses 24 clips to cover all combinations of four emotions, three speeds, and two styles. However, you could modify build_reference_library.py to generate a subset if your application only requires specific prosody variations, though this reduces the granularity of control.

How does the markup parser handle nested or conflicting tags?

The parser in markup.py processes tags sequentially, with each state tag overwriting the previous value for that dimension. If you specify [EMO:happy] followed by [EMO:frustrated], the latter takes effect immediately. Inline tags like [THINKING] create discrete silence segments rather than persisting state.

Can I use this implementation with voices other than Fish Audio S1?

The current architecture in chapter6/controllable-tts/tts.py is tightly coupled to the Fish Audio SDK, specifically the voice cloning API. To use alternative TTS engines, you would need to reimplement the synth_speech function to map the 24 reference profiles to your provider's API while maintaining the segment-based concatenation logic.

What is the latency impact of the 24-reference approach versus direct synthesis?

Using the full library requires sequential API calls for each prosody segment and subsequent FFmpeg concatenation. For low-latency applications, consider caching frequently used reference clips or using synth_direct_reference with a single neutral reference, accepting reduced prosodic control for faster response times.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →