# How to Configure OpenAI TTS for Garden Skills: Complete Setup Guide

> Learn how to configure OpenAI TTS for Garden Skills with our complete setup guide. Export your API key, set the TTS provider, and start synthesizing audio effortlessly.

- Repository: [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills)
- Tags: how-to-guide
- Published: 2026-08-29

---

**Configure the OpenAI TTS provider in Garden Skills by exporting `OPENAI_API_KEY` and setting `PRESENTATION_TTS=openai` before running the audio synthesis command, with optional environment variables for model selection and base URL customization.**

Garden Skills is an open-source framework for automated presentation generation that ships with a pluggable text-to-speech subsystem. By configuring the OpenAI TTS provider, you can generate high-quality narration audio for any skill that implements the `tts-providers` contract, including the Web Video Presentation skill.

## Understanding the TTS Provider Architecture

The Garden Skills framework defines a strict provider contract located in `skills/<skill-name>/templates/scripts/tts-providers/`. Every TTS backend must be a single executable shell script that exposes three mandatory functions: `tts_check`, `tts_install_help`, and `tts_synthesize`.

The runtime orchestration happens in [`skills/web-video-presentation/templates/scripts/synthesize-audio.sh`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/synthesize-audio.sh). This runner reads the `PRESENTATION_TTS` environment variable to select the provider, sources the corresponding script, and invokes `tts_synthesize` for every narration segment in your presentation.

## Provider Location and Dependencies

The OpenAI provider implementation resides at:

```

skills/web-video-presentation/templates/scripts/tts-providers/openai.sh

```

Before configuration, ensure your system has the required dependencies. The `tts_check` function in [`openai.sh`](https://github.com/ConardLi/garden-skills/blob/main/openai.sh) validates the presence of `curl` and `jq`. Install these via your package manager if missing:

```bash

# macOS

brew install curl jq

# Ubuntu/Debian

sudo apt-get install curl jq

```

## Step-by-Step OpenAI TTS Configuration

### Set the API Key

Export your OpenAI API key as an environment variable. The [`openai.sh`](https://github.com/ConardLi/garden-skills/blob/main/openai.sh) script references `$OPENAI_API_KEY` to authenticate requests to the Audio Speech API:

```bash
export OPENAI_API_KEY=sk-your-secret-key-here

```

### Select the Provider

Set `PRESENTATION_TTS` to `openai` to instruct the synthesis runner to use the OpenAI provider:

```bash
export PRESENTATION_TTS=openai

```

### Configure Model and Voice (Optional)

Customize the synthesis quality by setting `OPENAI_TTS_MODEL`. The default is `tts-1` (fast, low-cost). For higher fidelity audio, specify `tts-1-hd`:

```bash
export OPENAI_TTS_MODEL=tts-1-hd

```

Voice selection happens per narration segment. If not specified in the segment metadata, the provider defaults to `alloy`. Valid OpenAI voices include `alloy`, `echo`, `fable`, `onyx`, `nova`, and `shimmer`.

### Override the Base URL (Optional)

For corporate proxies, Azure OpenAI, or self-hosted compatible gateways, override the endpoint:

```bash
export OPENAI_BASE_URL=https://your-proxy.example.com/v1

```

If unset, the provider defaults to `https://api.openai.com/v1`.

## How the OpenAI Provider Implements Synthesis

Inside [`openai.sh`](https://github.com/ConardLi/garden-skills/blob/main/openai.sh), the `tts_synthesize` function constructs a JSON payload and streams the audio via curl. The implementation follows this pattern:

```bash
tts_synthesize() {
  local text="$1"
  local voice="$2"
  local out="$3"
  local base="${OPENAI_BASE_URL:-https://api.openai.com/v1}"
  local model="${OPENAI_TTS_MODEL:-tts-1}"
  
  local payload
  payload=$(jq -n \
    --arg t "$text" \
    --arg v "$voice" \
    --arg m "$model" \
    '{model:$m, input:$t, voice:$v, response_format:"mp3"}')
    
  curl -fsS -o "$out" -X POST "$base/audio/speech" \
    -H "Authorization: Bearer $OPENAI_API_KEY" \
    -H "Content-Type: application/json" \
    -d "$payload"
}

```

The script uses `jq` to safely escape JSON parameters, preventing shell injection or quoting bugs. The `curl` command includes `-f` to fail on HTTP errors and `-sS` to suppress progress output while showing errors.

## Executing Audio Generation

With environment variables configured, run the synthesis command from your skill directory:

```bash
npm run synthesize-audio

```

The [`synthesize-audio.sh`](https://github.com/ConardLi/garden-skills/blob/main/synthesize-audio.sh) runner iterates through your presentation's narration segments, calls `tts_synthesize` for each, and writes MP3 files to the presentation's `audio/` directory or the specified output path.

## Troubleshooting Common Issues

- **Missing dependencies**: If `tts_check` reports missing tools, install `curl` and `jq` as shown in the prerequisites section.

- **Authentication failures**: Verify that `OPENAI_API_KEY` is exported in the same shell session running the synthesis command. The provider passes this directly to the `Authorization: Bearer` header.

- **Rate limiting**: The OpenAI Speech endpoint bills per second of generated audio. Monitor your usage in the OpenAI dashboard if you encounter quota errors.

- **Proxy configuration**: When using `OPENAI_BASE_URL`, ensure your proxy forwards the `Authorization` header unchanged, as [`openai.sh`](https://github.com/ConardLi/garden-skills/blob/main/openai.sh) does not implement additional authentication mechanisms.

## Summary

- Export `OPENAI_API_KEY` to authenticate with the OpenAI Audio Speech API.
- Set `PRESENTATION_TTS=openai` to select the provider in [`synthesize-audio.sh`](https://github.com/ConardLi/garden-skills/blob/main/synthesize-audio.sh).
- Optionally configure `OPENAI_TTS_MODEL` (default: `tts-1`) and `OPENAI_BASE_URL` for custom endpoints.
- The provider implements `tts_check`, `tts_install_help`, and `tts_synthesize` functions as defined in the Garden Skills contract.
- Generated audio is saved as MP3 files via direct streaming curl requests.

## Frequently Asked Questions

### What environment variables are required to configure OpenAI TTS in Garden Skills?

Only `OPENAI_API_KEY` and `PRESENTATION_TTS` are required. The [`openai.sh`](https://github.com/ConardLi/garden-skills/blob/main/openai.sh) provider reads `OPENAI_API_KEY` for authentication and the runner uses `PRESENTATION_TTS=openai` to locate the correct provider script. All other variables like `OPENAI_TTS_MODEL` and `OPENAI_BASE_URL` are optional and fallback to sensible defaults.

### How do I switch between different TTS voices when using the OpenAI provider?

Voice selection is typically passed per narration segment by the calling skill. If a segment does not specify a voice, the [`openai.sh`](https://github.com/ConardLi/garden-skills/blob/main/openai.sh) script defaults to `alloy`. You can explicitly request voices like `nova`, `shimmer`, or `onyx` in your presentation metadata, and the provider will include this value in the JSON payload sent to the OpenAI API.

### Can I use a self-hosted or proxy endpoint instead of the official OpenAI API?

Yes. Set `OPENAI_BASE_URL` to your proxy or Azure OpenAI endpoint before running synthesis. The provider constructs the final URL as `${OPENAI_BASE_URL}/audio/speech`, allowing you to route requests through corporate gateways or compatible self-hosted services. Ensure the proxy forwards the `Authorization` header containing your API key.

### What audio format does the OpenAI TTS provider output?

The provider always requests `response_format: "mp3"` from the OpenAI API and saves the binary response directly to the output file path specified by the runner. The resulting files are MP3 encoded at the bitrate and sample rate returned by OpenAI's Audio Speech endpoint, typically ready for immediate use in web presentations or video editing workflows.