# What TTS Providers Are Supported by the web-video-presentation Skill?

> Discover the TTS providers supported by the web-video-presentation skill. Explore options like MiniMax, OpenAI, ElevenLabs, and more, plus custom provider integration.

- Repository: [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills)
- Tags: api-reference
- Published: 2026-09-01

---

**The web-video-presentation skill supports seven TTS providers including MiniMax (default), OpenAI, ElevenLabs, edge-tts, macOS say, Azure Speech, and Google Cloud TTS, with a pluggable architecture that allows custom providers via shell scripts exposing a `tts_synthesize` function.**

The **web-video-presentation** skill in the `ConardLi/garden-skills` repository implements a modular text-to-speech architecture designed for flexibility. Rather than hardcoding vendor APIs, the skill loads provider implementations dynamically from `skills/web-video-presentation/templates/scripts/tts-providers/` based on runtime configuration. This design allows developers to swap between cloud services, local engines, or custom implementations without modifying the core skill logic.

## Built-in TTS Providers

The skill ships with two fully supported production providers and five additional sample implementations that demonstrate the provider contract.

### MiniMax (Default)

**MiniMax** serves as the default TTS provider when no explicit configuration is provided. The implementation resides in [`skills/web-video-presentation/templates/scripts/tts-providers/minimax.sh`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/minimax.sh) and interfaces with the MiniMax `mmx` CLI tool.

Authentication requires running `mmx auth login --api-key` before synthesis. The provider script implements the mandatory `tts_synthesize` function that accepts text input, output path, and voice parameters, converting them into MiniMax-specific API calls.

### OpenAI

The **OpenAI** provider at [`skills/web-video-presentation/templates/scripts/tts-providers/openai.sh`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/openai.sh) connects to either the standard OpenAI Audio Speech API or Azure OpenAI endpoints. This provider requires the `OPENAI_API_KEY` environment variable for authentication.

Unlike MiniMax, this implementation supports configurable base URLs, allowing enterprises to route requests through private Azure deployments while maintaining the same interface contract.

## Sample Provider Scripts

Beyond the two built-in providers, the repository includes ready-to-use sample scripts for alternative TTS backends. These follow the same three-function contract but may require additional setup.

### ElevenLabs

The **ElevenLabs** provider script ([`skills/web-video-presentation/templates/scripts/tts-providers/elevenlabs.sh`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/elevenlabs.sh)) supports high-quality multilingual voice synthesis. It requires the `ELEVENLABS_API_KEY` environment variable and supports voice ID selection through the `PRESENTATION_TTS_VOICE` variable.

### edge-tts

**edge-tts** provides a free, no-API-key alternative using Microsoft's Edge TTS service. The implementation in [`skills/web-video-presentation/templates/scripts/tts-providers/edge-tts.sh`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/edge-tts.sh) requires the Python package (`pip install edge-tts`) but offers zero-cost synthesis with decent quality for development and testing.

### macOS say

For offline usage on Apple hardware, the **macOS say** provider ([`skills/web-video-presentation/templates/scripts/tts-providers/say.sh`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/say.sh)) wraps the native `say` command. This provider requires `ffmpeg` to convert the generated `.aiff` files to `.mp3` format for web compatibility, making it ideal for local development without network dependencies.

### Azure Speech

The **Azure Speech** provider at [`skills/web-video-presentation/templates/scripts/tts-providers/azure.sh`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/azure.sh) integrates with Azure's Speech Service. Configuration requires both `AZURE_SPEECH_KEY` and `AZURE_SPEECH_REGION` environment variables. This provider suits enterprise deployments already invested in Microsoft's cognitive services infrastructure.

### Google Cloud TTS

The **Google Cloud TTS** implementation ([`skills/web-video-presentation/templates/scripts/tts-providers/gcloud.sh`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/gcloud.sh)) leverages Google Cloud's Text-to-Speech API. It requires the `gcloud` SDK installation and proper service account authentication, offering access to WaveNet and Neural2 voice models.

## How the Provider Architecture Works

The skill's runner script, [`skills/web-video-presentation/templates/scripts/synthesize-audio.sh`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/synthesize-audio.sh), implements a dynamic loading mechanism. At runtime, it reads the `PRESENTATION_TTS` environment variable and sources the corresponding script from `skills/web-video-presentation/templates/scripts/tts-providers/<name>.sh`.

According to the provider contract documented in [`skills/web-video-presentation/templates/scripts/tts-providers/README.md`](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/README.md), every provider script must expose:

- **`tts_synthesize`** (required): The core function accepting text, output path, and voice parameters
- **`tts_check`** (optional): Validation function verifying prerequisites and API connectivity
- **`tts_install_help`** (optional): Function displaying setup instructions when dependencies are missing

This architecture decouples the orchestration logic from vendor-specific implementations, allowing the skill to treat all TTS services as interchangeable black boxes.

## Switching Between Providers

Changing TTS providers requires no code changes—only environment variable modifications.

Use the default MiniMax provider without additional configuration:

```bash
npm run synthesize-audio

```

Switch to OpenAI by setting the provider variable:

```bash
PRESENTATION_TTS=openai npm run synthesize-audio

```

Configure ElevenLabs with a specific voice ID:

```bash
PRESENTATION_TTS=elevenlabs PRESENTATION_TTS_VOICE=21m00Tcm4TlvDq8ikWAM \
  npm run synthesize-audio

```

For debugging individual providers without the full pipeline, source the script directly and invoke its functions:

```bash
source skills/web-video-presentation/templates/scripts/tts-providers/edge-tts.sh
tts_check && tts_synthesize "测试一下" /tmp/test.mp3 "zh-CN-YunxiNeural"

```

## Creating Custom TTS Providers

Developers can extend the skill by creating new provider scripts in `skills/web-video-presentation/templates/scripts/tts-providers/`. A valid provider requires only a shell script named `<provider-name>.sh` implementing the `tts_synthesize` function with this signature:

```bash
tts_synthesize() {
  local text="$1"
  local output_file="$2"
  local voice="$3"
  # Implementation-specific logic here

}

```

Optional helper functions `tts_check` and `tts_install_help` improve user experience by validating prerequisites and displaying setup guidance. Once the file is placed in the providers directory, set `PRESENTATION_TTS=<provider-name>` to activate it.

## Summary

- The **web-video-presentation** skill supports **MiniMax** (default) and **OpenAI** as production-ready providers, plus **ElevenLabs**, **edge-tts**, **macOS say**, **Azure Speech**, and **Google Cloud TTS** as sample implementations.
- The **pluggable architecture** uses the `PRESENTATION_TTS` environment variable to dynamically load provider scripts from `skills/web-video-presentation/templates/scripts/tts-providers/`.
- Each provider must implement **`tts_synthesize`** and may optionally implement `tts_check` and `tts_install_help`.
- Switching providers requires no code changes—only environment variable configuration.
- Custom providers follow a simple three-function shell script contract defined in the README.

## Frequently Asked Questions

### How do I add a new TTS provider to the web-video-presentation skill?

Create a shell script in `skills/web-video-presentation/templates/scripts/tts-providers/` named `<your-provider>.sh` that implements the `tts_synthesize` function accepting text, output path, and voice parameters. Optionally add `tts_check` for validation and `tts_install_help` for setup instructions. Set `PRESENTATION_TTS=your-provider` to activate it.

### Which TTS provider works without an API key?

The **edge-tts** provider requires no API key and uses Microsoft's free Edge TTS service. The **macOS say** provider also requires no cloud credentials but only works on macOS systems and needs `ffmpeg` for audio format conversion.

### Can I use Azure OpenAI instead of the standard OpenAI endpoint?

Yes. The [`openai.sh`](https://github.com/ConardLi/garden-skills/blob/main/openai.sh) provider supports Azure OpenAI endpoints. Configure the `OPENAI_API_KEY` environment variable and ensure your Azure deployment details are properly set in the provider configuration or through additional environment variables handled by the OpenAI client.

### What happens if the TTS provider prerequisites are missing?

If a provider implements the optional `tts_check` function, the runner script will detect missing dependencies before attempting synthesis. If `tts_install_help` is implemented, the skill will display platform-specific setup instructions. Without these helpers, the script will fail when attempting to execute the vendor's CLI tool or API client.