How to Configure OpenAI TTS for the Web-Video-Presentation Skill

Configure OpenAI TTS for the web-video-presentation skill by exporting the OPENAI_API_KEY environment variable, ensuring the openai.sh provider script is executable, and launching the skill with the --tts-provider openai runtime flag.

The ConardLi/garden-skills repository ships the web-video-presentation skill with a provider-agnostic audio runner that supports multiple Text-to-Speech backends. By default, the skill includes both a MiniMax CLI provider and an OpenAI provider, allowing you to generate high-quality speech synthesis for interactive web-video demos using OpenAI's Audio API.

Setting Up Environment Credentials

The OpenAI TTS provider relies on environment variables for authentication and endpoint configuration. These values are read at runtime by the shell script located at skills/web-video-presentation/templates/scripts/tts-providers/openai.sh.

Exporting the OpenAI API Key

Set the OPENAI_API_KEY variable in your shell environment before running the skill. The provider script passes this key directly to the OpenAI API via curl.

export OPENAI_API_KEY=sk-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx

Never commit this key to version control. For production deployments, use a secret manager or your orchestration platform's environment variable injection (e.g., Docker secrets, Kubernetes env vars, or GitHub Actions encrypted secrets).

Configuring a Custom Base URL (Optional)

If you are using Azure OpenAI, a reverse proxy, or a compatible gateway, override the default endpoint by setting OPENAI_BASE_URL. The default value is https://api.openai.com/v1.

export OPENAI_BASE_URL=https://your-proxy-or-azure-openai.openai.azure.com/openai/deployments/your-deployment/v1

Locating the Provider Script

The OpenAI TTS implementation resides in the skills directory tree as a standalone shell script. Verify the file exists at:


skills/web-video-presentation/templates/scripts/tts-providers/openai.sh

This script wraps a curl command that sends JSON payloads to the OpenAI Audio Speech API. The skill automatically discovers all executable .sh files placed in the tts-providers/ directory, requiring no additional registration code to enable the provider.

Testing the TTS Provider Locally

Before integrating with the full skill pipeline, validate the provider independently. The openai.sh script accepts two arguments: a JSON payload file and an output audio file path.

First, create a payload file specifying the model, input text, voice, and output format:

cat > payload.json <<'EOF'
{
  "model": "tts-1",
  "input": "你好,欢迎使用 Garden 的 Web 视频演示技能。",
  "voice": "alloy",
  "output_format": "mp3"
}
EOF

Then invoke the provider script directly:

chmod +x skills/web-video-presentation/templates/scripts/tts-providers/openai.sh
./skills/web-video-presentation/templates/scripts/tts-providers/openai.sh payload.json output.mp3

If the request succeeds, output.mp3 contains the synthesized speech. If the API key is missing or invalid, the script exits with an error code, triggering the skill's fallback mechanism.

Enabling OpenAI TTS in the Skill

Once environment variables are configured, you must ensure the provider script is executable and select it at runtime.

Setting Execution Permissions

The skill only discovers providers that have executable permissions. Run the following command from the repository root:

chmod +x skills/web-video-presentation/templates/scripts/tts-providers/openai.sh

Without this step, the skill ignores the openai.sh file and defaults to the next available provider (MiniMax).

Runtime Provider Selection

Launch the skill with the explicit provider flag to use OpenAI TTS:

garden skill run web-video-presentation --tts-provider openai

If you omit the --tts-provider flag, the skill iterates through available providers in the tts-providers/ directory and selects the first executable one. If the OpenAI provider fails due to missing credentials, the skill automatically falls back to the MiniMax mmx-cli provider, ensuring the pipeline remains operational.

Understanding the Provider Architecture

According to the ConardLi/garden-skills source code, the web-video-presentation skill implements a pluggable TTS architecture defined in the top-level README.md and SKILL.md files. The runner scans skills/web-video-presentation/templates/scripts/tts-providers/ for executable shell scripts, treating each as a standalone TTS adapter. This design allows you to add custom providers by placing new executable scripts in the same directory without modifying the skill's core logic.

The OpenAI provider specifically handles the following parameters from the JSON payload:

  • model: The TTS model ID (e.g., tts-1, tts-1-hd)
  • input: The text to synthesize (maximum length enforced by OpenAI's API limits)
  • voice: The voice style (e.g., alloy, echo, fable, onyx, nova, shimmer)
  • output_format: The audio encoding format (e.g., mp3, opus, aac, flac, wav, pcm)

Summary

Frequently Asked Questions

What audio formats does the OpenAI TTS provider support?

The provider supports all formats specified in the OpenAI Audio API payload. Valid output_format values include mp3, opus, aac, flac, wav, and pcm. You define the desired format in the JSON payload passed to the openai.sh script, as documented in skills/web-video-presentation/templates/scripts/tts-providers/README.md.

How does the skill handle missing or invalid OpenAI API keys?

If the OPENAI_API_KEY environment variable is unset or the API returns an authentication error, the openai.sh script exits with a non-zero status. The skill's provider-agnostic runner catches this failure and automatically falls back to the MiniMax mmx-cli provider, assuming it is installed and configured. This ensures the video generation pipeline continues to function even if one provider is unavailable.

Can I use Azure OpenAI or a proxy instead of the official API endpoint?

Yes. Set the OPENAI_BASE_URL environment variable to your Azure OpenAI endpoint or proxy gateway before running the skill. The shell script uses this variable to construct the request URL, replacing the default https://api.openai.com/v1 base path. This allows integration with private deployments or API gateways without modifying the provider script's source code.

Where are the TTS provider scripts located in the repository?

The built-in TTS providers reside in skills/web-video-presentation/templates/scripts/tts-providers/. The OpenAI provider specifically lives at skills/web-video-presentation/templates/scripts/tts-providers/openai.sh. The skill discovers providers by scanning this directory for executable files, as detailed in the SKILL.md manifest and the top-level README.md under the "Pluggable TTS" section.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →