How to Integrate Custom TTS Providers in Garden Skills: A Complete Developer Guide
Garden Skills synthesizes slide narration through a provider-agnostic shell runner that dynamically loads TTS implementations from the tts-providers directory, requiring only a tts_synthesize function to convert text to audio.
Garden Skills supports custom text-to-speech (TTS) integration through a flexible plugin architecture defined in skills/web-video-presentation/templates/scripts/tts-providers/README.md. The system uses a shell-based abstraction layer that delegates audio generation to external scripts, allowing you to integrate proprietary APIs, local engines, or cloud services without modifying the core synthesis runner. By implementing a minimal function contract, your custom provider can process narration segments from audio-segments.json and output web-playable audio files.
The Provider Contract
Each provider script must export specific functions that the runner consumes during the synthesis pipeline. The contract consists of one mandatory and two optional functions:
-
tts_synthesize <text> <out_path> [<voice]— The required core function that converts UTF-8 text into an audio file (MP3 by default) at the specified output path. The optional voice parameter allows users to specify different speakers or styles. -
tts_check— An optional sanity check executed once before the first segment. Use this to verify CLI tool availability, API key presence, or network connectivity. Return non-zero to signal configuration errors. -
tts_install_help— An optional helper invoked whentts_checkfails. This function prints setup instructions to stderr, guiding users through API key configuration or dependency installation.
The runner iterates over audio-segments.json, skips already-generated files, and invokes tts_synthesize for each pending segment. Non-zero exit codes mark segments as FAILED but allow the pipeline to continue processing remaining items.
Step-by-Step Integration Guide
Follow these steps to add a custom TTS provider to Garden Skills:
-
Create a new script in
skills/web-video-presentation/templates/scripts/tts-providers/with a descriptive name (e.g.,my-tts.sh). -
Implement the required contract by defining
tts_synthesizewith the exact signaturetts_synthesize <text> <out_path> [<voice]. The function must write a valid audio file toout_path. -
Add optional lifecycle functions (
tts_checkandtts_install_help) to provide validation and user guidance during setup. -
Make the script executable using
chmod +x skills/web-video-presentation/templates/scripts/tts-providers/my-tts.sh. -
Activate the provider via environment variable or CLI flag when running the synthesis command.
Minimal Implementation Example
Below is a complete skeleton for a REST-based TTS provider. This example demonstrates API key validation, error handling, and the required function signatures as specified in the Garden Skills source:
#!/usr/bin/env bash
# File: skills/web-video-presentation/templates/scripts/tts-providers/my-tts.sh
# Optional: verify that required CLI tools and API keys are present
tts_check() {
command -v curl >/dev/null || { echo "✗ curl not found" >&2; return 1; }
[[ -n "${MY_TTS_API_KEY:-}" ]] || { echo "✗ MY_TTS_API_KEY not set" >&2; return 1; }
}
# Optional: provide setup guidance when checks fail
tts_install_help() {
cat <<'EOF' >&2
Set up your My-TTS provider:
export MY_TTS_API_KEY=your_key_here # get a key from https://example.com
EOF
}
# Required: synthesize text to the specified output file
tts_synthesize() {
local text="$1"
local out="$2"
local voice="${3:-default-voice}"
# Build JSON payload safely using jq
local payload
payload=$(jq -n --arg t "$text" --arg v "$voice" '{text:$t, voice:$v}')
curl -fsS -o "$out" -X POST "https://api.my-tts.com/v1/speak" \
-H "Authorization: Bearer $MY_TTS_API_KEY" \
-H "Content-Type: application/json" \
-d "$payload"
}
Activating Your Custom Provider
Garden Skills discovers providers by matching the script name against the PRESENTATION_TTS value. You can specify your provider using either environment variables or command-line flags:
# Method 1: Environment variable
PRESENTATION_TTS=my-tts npm run synthesize-audio
# Method 2: CLI flag
npm run synthesize-audio -- --provider=my-tts
The synthesize-audio.sh runner automatically resolves the script path, executes tts_check once for validation, then processes each narration segment through your tts_synthesize implementation.
Reference Implementations
Study these existing providers in the Garden Skills repository to understand different architectural patterns:
openai.sh— Demonstrates REST API integration usingcurlandjqfor JSON processingminimax.sh— Shows CLI-based provider implementation for local tool executionedge-tts.sh— Example of a provider requiring no API keys (free service integration)say.sh— macOS-specific offline fallback using the built-insaycommand
These files reside in skills/web-video-presentation/templates/scripts/tts-providers/ and serve as authoritative references for error handling, voice parameter usage, and audio format compliance.
Summary
- Garden Skills uses a shell-based provider system located in
skills/web-video-presentation/templates/scripts/tts-providers/ - Custom providers must implement
tts_synthesize; optionally addtts_checkandtts_install_helpfor better UX - Activation occurs via the
PRESENTATION_TTSenvironment variable or--providerCLI flag - The runner processes
audio-segments.jsonsequentially, marking failed segments but continuing execution - Reference implementations including
openai.shandedge-tts.shdemonstrate both API and CLI integration patterns
Frequently Asked Questions
What audio formats must custom TTS providers output?
Custom providers should generate MP3 files by default, as this format offers universal browser compatibility and efficient compression. The tts_synthesize function receives an output path parameter—simply write any web-playable audio format (MP3, WAV, OGG) to that location. The Garden Skills player detects the file extension automatically, though MP3 remains the tested standard across built-in providers like openai.sh and minimax.sh.
How do I debug a provider that fails during synthesis?
First, run tts_check manually in your terminal to verify API keys and dependencies. If the check passes but synthesis fails, examine the exit codes: the runner marks segments as FAILED when tts_synthesize returns non-zero, but continues processing. Check stderr output for error messages from your API calls. You can also test your script in isolation by calling tts_synthesize "test text" /tmp/test.mp3 directly before invoking the full pipeline.
Can I create a provider that works offline without API keys?
Yes, local TTS engines are fully supported. The say.sh provider demonstrates offline macOS integration using the built-in say command with no external dependencies. For Linux or cross-platform offline support, implement tts_synthesize to call local binaries like espeak, piper, or coqui-tts. Ensure these tools are available in the system PATH or use tts_check to verify their installation before synthesis begins.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →