How to Integrate Custom TTS Providers in Garden Skills: A Complete Developer Guide

Garden Skills synthesizes slide narration through a provider-agnostic shell runner that dynamically loads TTS implementations from the tts-providers directory, requiring only a tts_synthesize function to convert text to audio.

Garden Skills supports custom text-to-speech (TTS) integration through a flexible plugin architecture defined in skills/web-video-presentation/templates/scripts/tts-providers/README.md. The system uses a shell-based abstraction layer that delegates audio generation to external scripts, allowing you to integrate proprietary APIs, local engines, or cloud services without modifying the core synthesis runner. By implementing a minimal function contract, your custom provider can process narration segments from audio-segments.json and output web-playable audio files.

The Provider Contract

Each provider script must export specific functions that the runner consumes during the synthesis pipeline. The contract consists of one mandatory and two optional functions:

  • tts_synthesize <text> <out_path> [<voice] — The required core function that converts UTF-8 text into an audio file (MP3 by default) at the specified output path. The optional voice parameter allows users to specify different speakers or styles.

  • tts_check — An optional sanity check executed once before the first segment. Use this to verify CLI tool availability, API key presence, or network connectivity. Return non-zero to signal configuration errors.

  • tts_install_help — An optional helper invoked when tts_check fails. This function prints setup instructions to stderr, guiding users through API key configuration or dependency installation.

The runner iterates over audio-segments.json, skips already-generated files, and invokes tts_synthesize for each pending segment. Non-zero exit codes mark segments as FAILED but allow the pipeline to continue processing remaining items.

Step-by-Step Integration Guide

Follow these steps to add a custom TTS provider to Garden Skills:

  1. Create a new script in skills/web-video-presentation/templates/scripts/tts-providers/ with a descriptive name (e.g., my-tts.sh).

  2. Implement the required contract by defining tts_synthesize with the exact signature tts_synthesize <text> <out_path> [<voice]. The function must write a valid audio file to out_path.

  3. Add optional lifecycle functions (tts_check and tts_install_help) to provide validation and user guidance during setup.

  4. Make the script executable using chmod +x skills/web-video-presentation/templates/scripts/tts-providers/my-tts.sh.

  5. Activate the provider via environment variable or CLI flag when running the synthesis command.

Minimal Implementation Example

Below is a complete skeleton for a REST-based TTS provider. This example demonstrates API key validation, error handling, and the required function signatures as specified in the Garden Skills source:

#!/usr/bin/env bash

# File: skills/web-video-presentation/templates/scripts/tts-providers/my-tts.sh

# Optional: verify that required CLI tools and API keys are present

tts_check() {
  command -v curl >/dev/null || { echo "✗ curl not found" >&2; return 1; }
  [[ -n "${MY_TTS_API_KEY:-}" ]] || { echo "✗ MY_TTS_API_KEY not set" >&2; return 1; }
}

# Optional: provide setup guidance when checks fail

tts_install_help() {
  cat <<'EOF' >&2
Set up your My-TTS provider:
  export MY_TTS_API_KEY=your_key_here   # get a key from https://example.com

EOF
}

# Required: synthesize text to the specified output file

tts_synthesize() {
  local text="$1"
  local out="$2"
  local voice="${3:-default-voice}"

  # Build JSON payload safely using jq

  local payload
  payload=$(jq -n --arg t "$text" --arg v "$voice" '{text:$t, voice:$v}')
  
  curl -fsS -o "$out" -X POST "https://api.my-tts.com/v1/speak" \
    -H "Authorization: Bearer $MY_TTS_API_KEY" \
    -H "Content-Type: application/json" \
    -d "$payload"
}

Activating Your Custom Provider

Garden Skills discovers providers by matching the script name against the PRESENTATION_TTS value. You can specify your provider using either environment variables or command-line flags:


# Method 1: Environment variable

PRESENTATION_TTS=my-tts npm run synthesize-audio

# Method 2: CLI flag

npm run synthesize-audio -- --provider=my-tts

The synthesize-audio.sh runner automatically resolves the script path, executes tts_check once for validation, then processes each narration segment through your tts_synthesize implementation.

Reference Implementations

Study these existing providers in the Garden Skills repository to understand different architectural patterns:

  • openai.sh — Demonstrates REST API integration using curl and jq for JSON processing
  • minimax.sh — Shows CLI-based provider implementation for local tool execution
  • edge-tts.sh — Example of a provider requiring no API keys (free service integration)
  • say.sh — macOS-specific offline fallback using the built-in say command

These files reside in skills/web-video-presentation/templates/scripts/tts-providers/ and serve as authoritative references for error handling, voice parameter usage, and audio format compliance.

Summary

  • Garden Skills uses a shell-based provider system located in skills/web-video-presentation/templates/scripts/tts-providers/
  • Custom providers must implement tts_synthesize; optionally add tts_check and tts_install_help for better UX
  • Activation occurs via the PRESENTATION_TTS environment variable or --provider CLI flag
  • The runner processes audio-segments.json sequentially, marking failed segments but continuing execution
  • Reference implementations including openai.sh and edge-tts.sh demonstrate both API and CLI integration patterns

Frequently Asked Questions

What audio formats must custom TTS providers output?

Custom providers should generate MP3 files by default, as this format offers universal browser compatibility and efficient compression. The tts_synthesize function receives an output path parameter—simply write any web-playable audio format (MP3, WAV, OGG) to that location. The Garden Skills player detects the file extension automatically, though MP3 remains the tested standard across built-in providers like openai.sh and minimax.sh.

How do I debug a provider that fails during synthesis?

First, run tts_check manually in your terminal to verify API keys and dependencies. If the check passes but synthesis fails, examine the exit codes: the runner marks segments as FAILED when tts_synthesize returns non-zero, but continues processing. Check stderr output for error messages from your API calls. You can also test your script in isolation by calling tts_synthesize "test text" /tmp/test.mp3 directly before invoking the full pipeline.

Can I create a provider that works offline without API keys?

Yes, local TTS engines are fully supported. The say.sh provider demonstrates offline macOS integration using the built-in say command with no external dependencies. For Linux or cross-platform offline support, implement tts_synthesize to call local binaries like espeak, piper, or coqui-tts. Ensure these tools are available in the system PATH or use tts_check to verify their installation before synthesis begins.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →