How to Extend the TTS Architecture with New Providers in Garden-Skills

Adding a new TTS provider requires implementing the TtsProvider interface in src/tts/providers/, registering your class in the provider registry, and updating tts-config.json to point to your implementation.

The Garden-Skills repository provides a modular text-to-speech pipeline that converts chapter narration scripts into audio files. This architecture uses a provider pattern that abstracts external TTS services behind a common TypeScript interface, making it trivial to swap between Google Cloud, AWS Polly, or custom implementations. Because the pipeline in extract-narrations.ts generates a flat JSON manifest consumed by the synthesis stage, you can extend the TTS architecture with new providers without modifying existing chapter code.

Understanding the TTS Provider Contract

The foundation of the extensible architecture is the TtsProvider interface defined in src/tts/provider.ts. This contract requires a single method:

interface TtsProvider {
  synthesize(text: string): Promise<Uint8Array>;
}

The synthesize method accepts a raw text string and returns a Promise resolving to a Uint8Array containing MP3 or OGG audio data. Existing implementations in src/tts/providers/google.ts and src/tts/providers/aws.ts demonstrate how to wrap external APIs to satisfy this contract.

Implementing a Custom TTS Provider

Follow these three steps to add support for a new speech synthesis service.

Step 1 – Create the Provider Class

Create a new TypeScript file in src/tts/providers/. Your class must implement the TtsProvider interface and handle authentication, API communication, and error handling. Here is a complete implementation for a hypothetical MySpeech service:

// src/tts/providers/my-speech.ts
import type { TtsProvider } from "../provider";

export class MySpeechProvider implements TtsProvider {
  private readonly apiKey: string;

  constructor(apiKey: string) {
    this.apiKey = apiKey;
  }

  async synthesize(text: string): Promise<Uint8Array> {
    const response = await fetch("https://api.myspeech.io/v1/synthesize", {
      method: "POST",
      headers: {
        "Content-Type": "application/json",
        "Authorization": `Bearer ${this.apiKey}`,
      },
      body: JSON.stringify({ text, voice: "en-US-Standard-A" }),
    });

    if (!response.ok) {
      const err = await response.text();
      throw new Error(`MySpeech failed: ${err}`);
    }

    const arrayBuffer = await response.arrayBuffer();
    return new Uint8Array(arrayBuffer);
  }
}

Ensure your implementation returns raw audio bytes that the pipeline can write directly to disk.

Step 2 – Register in the Provider Registry

The central registry at src/tts/providers/index.ts maps string keys to provider instances. Import your class and add it to the providers record:

// src/tts/providers/index.ts
import { MySpeechProvider } from "./my-speech";
import type { TtsProvider } from "../provider";
import { GoogleProvider } from "./google";
import { AwsPollyProvider } from "./aws";

const providers: Record<string, TtsProvider> = {
  google: new GoogleProvider(),
  aws: new AwsPollyProvider(),
  myspeech: new MySpeechProvider(process.env.MYSPEECH_API_KEY ?? ""),
};

export default providers;

The key you assign here (myspeech) becomes the identifier used in configuration files.

Step 3 – Update the Configuration

Modify tts-config.json in the repository root to specify your new provider as the default:

{
  "defaultProvider": "myspeech",
  "outputDir": "public/audio"
}

When you run npm run generate-tts, the pipeline reads this configuration, instantiates your provider class, and executes the synthesis workflow.

How the Pipeline Processes Audio Generation

Understanding the full workflow helps debug integration issues. The process flow in Garden-Skills operates in three distinct stages:

  1. Narration Extraction – The script skills/web-video-presentation/templates/scripts/extract-narrations.ts traverses the chapter registry, loads each chapter's narrations.ts file, and flattens the content into audio-segments.json. Each entry contains the chapter ID, step number, raw text, and target filename.

  2. Provider Selection – The build script imports the provider map from src/tts/providers/index.ts and selects the implementation matching the defaultProvider key in tts-config.json.

  3. Synthesis and Caching – For each segment in audio-segments.json, the pipeline checks if public/audio/<chapter>/<step>.mp3 already exists. If missing, it calls your provider's synthesize method and writes the resulting Uint8Array to disk. This caching layer makes incremental builds fast when updating individual chapters.

Summary

To extend the TTS architecture with new providers in Garden-Skills:

  • Implement the TtsProvider interface with a class that converts text to a Uint8Array of audio bytes.
  • Register your implementation in src/tts/providers/index.ts by adding it to the provider record with a unique string key.
  • Configure the pipeline by setting the defaultProvider field in tts-config.json to your registered key.

The modular design ensures that extract-narrations.ts and existing chapter definitions require zero modifications when adding new voice services.

Frequently Asked Questions

What interface must new TTS providers implement in Garden-Skills?

All providers must implement the TtsProvider interface defined in src/tts/provider.ts, which requires a single synthesize(text: string): Promise<Uint8Array> method. This method must return a Promise that resolves to raw audio data as a Uint8Array, typically containing MP3 or OGG formatted bytes.

Where should I place my custom TTS provider implementation?

Create your provider file inside the src/tts/providers/ directory, following the naming convention [service-name].ts. After implementation, register the class in src/tts/providers/index.ts by importing it and adding it to the providers record object that maps configuration keys to provider instances.

How does the pipeline handle audio file caching?

The pipeline automatically checks for existing files at public/audio/<chapter>/<step>.mp3 before invoking your provider's synthesize method. If the target file already exists, the step is skipped entirely. This incremental behavior prevents unnecessary API calls and speeds up regeneration when iterating on specific chapters.

Can I switch between multiple TTS providers without code changes?

Yes. Because providers are registered by key in src/tts/providers/index.ts, you can switch implementations by changing the defaultProvider value in tts-config.json. No modifications to the synthesis logic or chapter files are required, as the pipeline dynamically instantiates whichever provider matches the configuration key.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →