How to Configure OpenAI TTS for Garden Skills: Complete Setup Guide
Configure the OpenAI TTS provider in Garden Skills by exporting OPENAI_API_KEY and setting PRESENTATION_TTS=openai before running the audio synthesis command, with optional environment variables for model selection and base URL customization.
Garden Skills is an open-source framework for automated presentation generation that ships with a pluggable text-to-speech subsystem. By configuring the OpenAI TTS provider, you can generate high-quality narration audio for any skill that implements the tts-providers contract, including the Web Video Presentation skill.
Understanding the TTS Provider Architecture
The Garden Skills framework defines a strict provider contract located in skills/<skill-name>/templates/scripts/tts-providers/. Every TTS backend must be a single executable shell script that exposes three mandatory functions: tts_check, tts_install_help, and tts_synthesize.
The runtime orchestration happens in skills/web-video-presentation/templates/scripts/synthesize-audio.sh. This runner reads the PRESENTATION_TTS environment variable to select the provider, sources the corresponding script, and invokes tts_synthesize for every narration segment in your presentation.
Provider Location and Dependencies
The OpenAI provider implementation resides at:
skills/web-video-presentation/templates/scripts/tts-providers/openai.sh
Before configuration, ensure your system has the required dependencies. The tts_check function in openai.sh validates the presence of curl and jq. Install these via your package manager if missing:
# macOS
brew install curl jq
# Ubuntu/Debian
sudo apt-get install curl jq
Step-by-Step OpenAI TTS Configuration
Set the API Key
Export your OpenAI API key as an environment variable. The openai.sh script references $OPENAI_API_KEY to authenticate requests to the Audio Speech API:
export OPENAI_API_KEY=sk-your-secret-key-here
Select the Provider
Set PRESENTATION_TTS to openai to instruct the synthesis runner to use the OpenAI provider:
export PRESENTATION_TTS=openai
Configure Model and Voice (Optional)
Customize the synthesis quality by setting OPENAI_TTS_MODEL. The default is tts-1 (fast, low-cost). For higher fidelity audio, specify tts-1-hd:
export OPENAI_TTS_MODEL=tts-1-hd
Voice selection happens per narration segment. If not specified in the segment metadata, the provider defaults to alloy. Valid OpenAI voices include alloy, echo, fable, onyx, nova, and shimmer.
Override the Base URL (Optional)
For corporate proxies, Azure OpenAI, or self-hosted compatible gateways, override the endpoint:
export OPENAI_BASE_URL=https://your-proxy.example.com/v1
If unset, the provider defaults to https://api.openai.com/v1.
How the OpenAI Provider Implements Synthesis
Inside openai.sh, the tts_synthesize function constructs a JSON payload and streams the audio via curl. The implementation follows this pattern:
tts_synthesize() {
local text="$1"
local voice="$2"
local out="$3"
local base="${OPENAI_BASE_URL:-https://api.openai.com/v1}"
local model="${OPENAI_TTS_MODEL:-tts-1}"
local payload
payload=$(jq -n \
--arg t "$text" \
--arg v "$voice" \
--arg m "$model" \
'{model:$m, input:$t, voice:$v, response_format:"mp3"}')
curl -fsS -o "$out" -X POST "$base/audio/speech" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d "$payload"
}
The script uses jq to safely escape JSON parameters, preventing shell injection or quoting bugs. The curl command includes -f to fail on HTTP errors and -sS to suppress progress output while showing errors.
Executing Audio Generation
With environment variables configured, run the synthesis command from your skill directory:
npm run synthesize-audio
The synthesize-audio.sh runner iterates through your presentation's narration segments, calls tts_synthesize for each, and writes MP3 files to the presentation's audio/ directory or the specified output path.
Troubleshooting Common Issues
-
Missing dependencies: If
tts_checkreports missing tools, installcurlandjqas shown in the prerequisites section. -
Authentication failures: Verify that
OPENAI_API_KEYis exported in the same shell session running the synthesis command. The provider passes this directly to theAuthorization: Bearerheader. -
Rate limiting: The OpenAI Speech endpoint bills per second of generated audio. Monitor your usage in the OpenAI dashboard if you encounter quota errors.
-
Proxy configuration: When using
OPENAI_BASE_URL, ensure your proxy forwards theAuthorizationheader unchanged, asopenai.shdoes not implement additional authentication mechanisms.
Summary
- Export
OPENAI_API_KEYto authenticate with the OpenAI Audio Speech API. - Set
PRESENTATION_TTS=openaito select the provider insynthesize-audio.sh. - Optionally configure
OPENAI_TTS_MODEL(default:tts-1) andOPENAI_BASE_URLfor custom endpoints. - The provider implements
tts_check,tts_install_help, andtts_synthesizefunctions as defined in the Garden Skills contract. - Generated audio is saved as MP3 files via direct streaming curl requests.
Frequently Asked Questions
What environment variables are required to configure OpenAI TTS in Garden Skills?
Only OPENAI_API_KEY and PRESENTATION_TTS are required. The openai.sh provider reads OPENAI_API_KEY for authentication and the runner uses PRESENTATION_TTS=openai to locate the correct provider script. All other variables like OPENAI_TTS_MODEL and OPENAI_BASE_URL are optional and fallback to sensible defaults.
How do I switch between different TTS voices when using the OpenAI provider?
Voice selection is typically passed per narration segment by the calling skill. If a segment does not specify a voice, the openai.sh script defaults to alloy. You can explicitly request voices like nova, shimmer, or onyx in your presentation metadata, and the provider will include this value in the JSON payload sent to the OpenAI API.
Can I use a self-hosted or proxy endpoint instead of the official OpenAI API?
Yes. Set OPENAI_BASE_URL to your proxy or Azure OpenAI endpoint before running synthesis. The provider constructs the final URL as ${OPENAI_BASE_URL}/audio/speech, allowing you to route requests through corporate gateways or compatible self-hosted services. Ensure the proxy forwards the Authorization header containing your API key.
What audio format does the OpenAI TTS provider output?
The provider always requests response_format: "mp3" from the OpenAI API and saves the binary response directly to the output file path specified by the runner. The resulting files are MP3 encoded at the bitrate and sample rate returned by OpenAI's Audio Speech endpoint, typically ready for immediate use in web presentations or video editing workflows.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →