# How to Implement Controllable TTS for Agent Voice Output Using Fish Audio S1

> Implement controllable TTS for agent voice output with Fish Audio S1. Dynamically adjust emotion, speed, and style using voice clips and markup tags.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: how-to-guide
- Published: 2026-08-17

---

**Controllable TTS enables dynamic manipulation of emotional tone, speaking speed, and conversational style by selecting from a curated library of 24 reference voice clips and parsing inline markup tags.**

The `bojieli/ai-agent-book` repository provides a complete implementation of **controllable text-to-speech** built on Fish Audio S1. This architecture allows AI agents to express happiness, frustration, or thinking states while adjusting tempo and formality on demand. The system decomposes the synthesis pipeline into three specialized layers that work together to transform marked-up scripts into expressive audio.

## The Three-Layer Architecture

The implementation separates concerns into distinct modules that handle voice storage, script interpretation, and audio generation.

### Reference Library

The **voice library** stores 24 real-voice clips covering every combination of emotion, speed, and style. Located in [`chapter6/controllable-tts/voice_library.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter6/controllable-tts/voice_library.py), this module manages the manifest that maps **prosody profiles** to reference audio files. The library handles neutral, happy, frustrated, and thinking emotions crossed with slow, normal, and fast speeds, further divided into formal and casual styles.

### Markup Parser

The parser in [`chapter6/controllable-tts/markup.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter6/controllable-tts/markup.py) converts annotated scripts into structured segments. It processes **state tags** like `[EMO:happy]`, `[SPEED:fast]`, and `[STYLE:casual]` that persist until changed, and **inline tags** such as `[THINKING]`, `<pause>`, or `[BREATH]` that insert non-verbal cues or silences.

### Synthesis Engine

The [`tts.py`](https://github.com/bojieli/ai-agent-book/blob/main/tts.py) module orchestrates voice cloning via the Fish Audio SDK. For each speech segment, it selects the appropriate reference clip based on current markup state, calls `synth_speech` to generate audio, and concatenates results using FFmpeg. Silence segments are synthesized separately and merged into the final MP3.

## Setting Up the Reference Library

Before parsing scripts, you must build the 24-clip library that enables prosody control. This requires a Fish API key and a base reference ID representing your owned voice.

Run the build script from the repository root:

```bash
export FISH_API_KEY="your_api_key_here"
python -m chapter6.controllable-tts.build_reference_library \
    --base-reference-id <YOUR_BASE_REFERENCE_ID>

```

This creates a `reference_audio/` directory containing the clips and a [`manifest.json`](https://github.com/bojieli/ai-agent-book/blob/main/manifest.json) file that [`voice_library.py`](https://github.com/bojieli/ai-agent-book/blob/main/voice_library.py) uses to resolve prosody requests.

## Writing and Parsing Control Markup

Scripts use a lightweight markup syntax to specify voice characteristics. State tags modify all subsequent text until overridden, while inline tags create immediate effects.

```python
from chapter6.controllable_tts.markup import parse

demo_script = (
    "[EMO:happy][SPEED:fast][STYLE:casual]太好了！您的订单已确认。"
    "[THINKING]让我查一下发货时间。"
    "[EMO:neutral][SPEED:normal][STYLE:formal]预计明天下午送达。"
)

segments = parse(demo_script)
print(segments)   # List of Segment objects (speech or silence)

```

Supported tags include:
- **Emotion tags**: `[EMO:neutral]`, `[EMO:happy]`, `[EMO:frustrated]`, `[EMO:thinking]`
- **Speed tags**: `[SPEED:slow]`, `[SPEED:normal]`, `[SPEED:fast]`
- **Style tags**: `[STYLE:formal]`, `[STYLE:casual]`
- **Non-verbal markers**: `[THINKING]`, `[BREATH]`, `<pause>`

## Synthesizing Speech with Prosody Control

With segments parsed and the library built, the `synthesize_segments` function in [`chapter6/controllable-tts/tts.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter6/controllable-tts/tts.py) handles the full audio pipeline.

```python
from chapter6.controllable_tts.tts import synthesize_segments
from chapter6.controllable_tts.voice_library import load_voice_library, DEFAULT_MANIFEST

# Load the manifest created during the build step

library = load_voice_library(DEFAULT_MANIFEST)

# Generate controlled audio

output_path = "output/agent_output.mp3"
info = synthesize_segments(
    segments=segments,
    out_path=output_path,
    workdir="tmp_workdir",
    manifest_path=DEFAULT_MANIFEST,
)

print(f"Audio written to {output_path}")

```

For scenarios requiring only basic voice cloning without prosody shifting, use `synth_direct_reference`:

```python
from chapter6.controllable_tts.tts import synth_direct_reference

plain_text = "Hello, this is a plain TTS request."
synth_direct_reference(
    text=plain_text,
    reference_id=library["source_reference_id"],
    out_path="output/plain_tts.mp3",
)

```

## Running the Complete Demo

The repository includes [`demo.py`](https://github.com/bojieli/ai-agent-book/blob/main/demo.py) to compare synthesis methods side-by-side. It generates three variants of the same input to demonstrate the value of controllable TTS.

```bash
python -m chapter6.controllable_tts.demo \
    --text "[EMO:happy][SPEED:fast][STYLE:casual]你好！" \
    --output-dir ./demo_output

```

This produces:
- **A_no_control_markers.mp3**: Direct S1 synthesis without reference library control
- **B_single_reference.mp3**: Single-clip reference (neutral + normal + formal)
- **C_24_reference_library.mp3**: Fully controllable output using the complete 24-clip library

The command also writes an [`evidence.json`](https://github.com/bojieli/ai-agent-book/blob/main/evidence.json) file documenting the experimental configuration and segment metadata.

## Summary

- **Controllable TTS** requires a multi-clip reference library covering emotion, speed, and style dimensions.
- The [`voice_library.py`](https://github.com/bojieli/ai-agent-book/blob/main/voice_library.py) module manages 24 distinct reference clips indexed by prosody profile.
- [`markup.py`](https://github.com/bojieli/ai-agent-book/blob/main/markup.py) parses control tags into segments that drive reference selection during synthesis.
- [`tts.py`](https://github.com/bojieli/ai-agent-book/blob/main/tts.py) handles Fish Audio SDK integration, audio generation, and FFmpeg-based concatenation.
- The system supports both fully controlled synthesis via `synthesize_segments` and basic cloning via `synth_direct_reference`.

## Frequently Asked Questions

### What is the minimum number of reference clips needed for basic controllable TTS?

The implementation uses 24 clips to cover all combinations of four emotions, three speeds, and two styles. However, you could modify [`build_reference_library.py`](https://github.com/bojieli/ai-agent-book/blob/main/build_reference_library.py) to generate a subset if your application only requires specific prosody variations, though this reduces the granularity of control.

### How does the markup parser handle nested or conflicting tags?

The parser in [`markup.py`](https://github.com/bojieli/ai-agent-book/blob/main/markup.py) processes tags sequentially, with each state tag overwriting the previous value for that dimension. If you specify `[EMO:happy]` followed by `[EMO:frustrated]`, the latter takes effect immediately. Inline tags like `[THINKING]` create discrete silence segments rather than persisting state.

### Can I use this implementation with voices other than Fish Audio S1?

The current architecture in [`chapter6/controllable-tts/tts.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter6/controllable-tts/tts.py) is tightly coupled to the Fish Audio SDK, specifically the voice cloning API. To use alternative TTS engines, you would need to reimplement the `synth_speech` function to map the 24 reference profiles to your provider's API while maintaining the segment-based concatenation logic.

### What is the latency impact of the 24-reference approach versus direct synthesis?

Using the full library requires sequential API calls for each prosody segment and subsequent FFmpeg concatenation. For low-latency applications, consider caching frequently used reference clips or using `synth_direct_reference` with a single neutral reference, accepting reduced prosodic control for faster response times.