How to Configure OpenAI TTS for Garden Skills: Complete Setup Guide

Configure the OpenAI TTS provider in Garden Skills by exporting OPENAI_API_KEY and setting PRESENTATION_TTS=openai before running the audio synthesis command, with optional environment variables for model selection and base URL customization.

Garden Skills is an open-source framework for automated presentation generation that ships with a pluggable text-to-speech subsystem. By configuring the OpenAI TTS provider, you can generate high-quality narration audio for any skill that implements the tts-providers contract, including the Web Video Presentation skill.

Understanding the TTS Provider Architecture

The Garden Skills framework defines a strict provider contract located in skills/<skill-name>/templates/scripts/tts-providers/. Every TTS backend must be a single executable shell script that exposes three mandatory functions: tts_check, tts_install_help, and tts_synthesize.

The runtime orchestration happens in skills/web-video-presentation/templates/scripts/synthesize-audio.sh. This runner reads the PRESENTATION_TTS environment variable to select the provider, sources the corresponding script, and invokes tts_synthesize for every narration segment in your presentation.

Provider Location and Dependencies

The OpenAI provider implementation resides at:


skills/web-video-presentation/templates/scripts/tts-providers/openai.sh

Before configuration, ensure your system has the required dependencies. The tts_check function in openai.sh validates the presence of curl and jq. Install these via your package manager if missing:


# macOS

brew install curl jq

# Ubuntu/Debian

sudo apt-get install curl jq

Step-by-Step OpenAI TTS Configuration

Set the API Key

Export your OpenAI API key as an environment variable. The openai.sh script references $OPENAI_API_KEY to authenticate requests to the Audio Speech API:

export OPENAI_API_KEY=sk-your-secret-key-here

Select the Provider

Set PRESENTATION_TTS to openai to instruct the synthesis runner to use the OpenAI provider:

export PRESENTATION_TTS=openai

Configure Model and Voice (Optional)

Customize the synthesis quality by setting OPENAI_TTS_MODEL. The default is tts-1 (fast, low-cost). For higher fidelity audio, specify tts-1-hd:

export OPENAI_TTS_MODEL=tts-1-hd

Voice selection happens per narration segment. If not specified in the segment metadata, the provider defaults to alloy. Valid OpenAI voices include alloy, echo, fable, onyx, nova, and shimmer.

Override the Base URL (Optional)

For corporate proxies, Azure OpenAI, or self-hosted compatible gateways, override the endpoint:

export OPENAI_BASE_URL=https://your-proxy.example.com/v1

If unset, the provider defaults to https://api.openai.com/v1.

How the OpenAI Provider Implements Synthesis

Inside openai.sh, the tts_synthesize function constructs a JSON payload and streams the audio via curl. The implementation follows this pattern:

tts_synthesize() {
  local text="$1"
  local voice="$2"
  local out="$3"
  local base="${OPENAI_BASE_URL:-https://api.openai.com/v1}"
  local model="${OPENAI_TTS_MODEL:-tts-1}"
  
  local payload
  payload=$(jq -n \
    --arg t "$text" \
    --arg v "$voice" \
    --arg m "$model" \
    '{model:$m, input:$t, voice:$v, response_format:"mp3"}')
    
  curl -fsS -o "$out" -X POST "$base/audio/speech" \
    -H "Authorization: Bearer $OPENAI_API_KEY" \
    -H "Content-Type: application/json" \
    -d "$payload"
}

The script uses jq to safely escape JSON parameters, preventing shell injection or quoting bugs. The curl command includes -f to fail on HTTP errors and -sS to suppress progress output while showing errors.

Executing Audio Generation

With environment variables configured, run the synthesis command from your skill directory:

npm run synthesize-audio

The synthesize-audio.sh runner iterates through your presentation's narration segments, calls tts_synthesize for each, and writes MP3 files to the presentation's audio/ directory or the specified output path.

Troubleshooting Common Issues

  • Missing dependencies: If tts_check reports missing tools, install curl and jq as shown in the prerequisites section.

  • Authentication failures: Verify that OPENAI_API_KEY is exported in the same shell session running the synthesis command. The provider passes this directly to the Authorization: Bearer header.

  • Rate limiting: The OpenAI Speech endpoint bills per second of generated audio. Monitor your usage in the OpenAI dashboard if you encounter quota errors.

  • Proxy configuration: When using OPENAI_BASE_URL, ensure your proxy forwards the Authorization header unchanged, as openai.sh does not implement additional authentication mechanisms.

Summary

  • Export OPENAI_API_KEY to authenticate with the OpenAI Audio Speech API.
  • Set PRESENTATION_TTS=openai to select the provider in synthesize-audio.sh.
  • Optionally configure OPENAI_TTS_MODEL (default: tts-1) and OPENAI_BASE_URL for custom endpoints.
  • The provider implements tts_check, tts_install_help, and tts_synthesize functions as defined in the Garden Skills contract.
  • Generated audio is saved as MP3 files via direct streaming curl requests.

Frequently Asked Questions

What environment variables are required to configure OpenAI TTS in Garden Skills?

Only OPENAI_API_KEY and PRESENTATION_TTS are required. The openai.sh provider reads OPENAI_API_KEY for authentication and the runner uses PRESENTATION_TTS=openai to locate the correct provider script. All other variables like OPENAI_TTS_MODEL and OPENAI_BASE_URL are optional and fallback to sensible defaults.

How do I switch between different TTS voices when using the OpenAI provider?

Voice selection is typically passed per narration segment by the calling skill. If a segment does not specify a voice, the openai.sh script defaults to alloy. You can explicitly request voices like nova, shimmer, or onyx in your presentation metadata, and the provider will include this value in the JSON payload sent to the OpenAI API.

Can I use a self-hosted or proxy endpoint instead of the official OpenAI API?

Yes. Set OPENAI_BASE_URL to your proxy or Azure OpenAI endpoint before running synthesis. The provider constructs the final URL as ${OPENAI_BASE_URL}/audio/speech, allowing you to route requests through corporate gateways or compatible self-hosted services. Ensure the proxy forwards the Authorization header containing your API key.

What audio format does the OpenAI TTS provider output?

The provider always requests response_format: "mp3" from the OpenAI API and saves the binary response directly to the output file path specified by the runner. The resulting files are MP3 encoded at the bitrate and sample rate returned by OpenAI's Audio Speech endpoint, typically ready for immediate use in web presentations or video editing workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →