What TTS Providers Are Supported by the web-video-presentation Skill?
The web-video-presentation skill supports seven TTS providers including MiniMax (default), OpenAI, ElevenLabs, edge-tts, macOS say, Azure Speech, and Google Cloud TTS, with a pluggable architecture that allows custom providers via shell scripts exposing a tts_synthesize function.
The web-video-presentation skill in the ConardLi/garden-skills repository implements a modular text-to-speech architecture designed for flexibility. Rather than hardcoding vendor APIs, the skill loads provider implementations dynamically from skills/web-video-presentation/templates/scripts/tts-providers/ based on runtime configuration. This design allows developers to swap between cloud services, local engines, or custom implementations without modifying the core skill logic.
Built-in TTS Providers
The skill ships with two fully supported production providers and five additional sample implementations that demonstrate the provider contract.
MiniMax (Default)
MiniMax serves as the default TTS provider when no explicit configuration is provided. The implementation resides in skills/web-video-presentation/templates/scripts/tts-providers/minimax.sh and interfaces with the MiniMax mmx CLI tool.
Authentication requires running mmx auth login --api-key before synthesis. The provider script implements the mandatory tts_synthesize function that accepts text input, output path, and voice parameters, converting them into MiniMax-specific API calls.
OpenAI
The OpenAI provider at skills/web-video-presentation/templates/scripts/tts-providers/openai.sh connects to either the standard OpenAI Audio Speech API or Azure OpenAI endpoints. This provider requires the OPENAI_API_KEY environment variable for authentication.
Unlike MiniMax, this implementation supports configurable base URLs, allowing enterprises to route requests through private Azure deployments while maintaining the same interface contract.
Sample Provider Scripts
Beyond the two built-in providers, the repository includes ready-to-use sample scripts for alternative TTS backends. These follow the same three-function contract but may require additional setup.
ElevenLabs
The ElevenLabs provider script (skills/web-video-presentation/templates/scripts/tts-providers/elevenlabs.sh) supports high-quality multilingual voice synthesis. It requires the ELEVENLABS_API_KEY environment variable and supports voice ID selection through the PRESENTATION_TTS_VOICE variable.
edge-tts
edge-tts provides a free, no-API-key alternative using Microsoft's Edge TTS service. The implementation in skills/web-video-presentation/templates/scripts/tts-providers/edge-tts.sh requires the Python package (pip install edge-tts) but offers zero-cost synthesis with decent quality for development and testing.
macOS say
For offline usage on Apple hardware, the macOS say provider (skills/web-video-presentation/templates/scripts/tts-providers/say.sh) wraps the native say command. This provider requires ffmpeg to convert the generated .aiff files to .mp3 format for web compatibility, making it ideal for local development without network dependencies.
Azure Speech
The Azure Speech provider at skills/web-video-presentation/templates/scripts/tts-providers/azure.sh integrates with Azure's Speech Service. Configuration requires both AZURE_SPEECH_KEY and AZURE_SPEECH_REGION environment variables. This provider suits enterprise deployments already invested in Microsoft's cognitive services infrastructure.
Google Cloud TTS
The Google Cloud TTS implementation (skills/web-video-presentation/templates/scripts/tts-providers/gcloud.sh) leverages Google Cloud's Text-to-Speech API. It requires the gcloud SDK installation and proper service account authentication, offering access to WaveNet and Neural2 voice models.
How the Provider Architecture Works
The skill's runner script, skills/web-video-presentation/templates/scripts/synthesize-audio.sh, implements a dynamic loading mechanism. At runtime, it reads the PRESENTATION_TTS environment variable and sources the corresponding script from skills/web-video-presentation/templates/scripts/tts-providers/<name>.sh.
According to the provider contract documented in skills/web-video-presentation/templates/scripts/tts-providers/README.md, every provider script must expose:
tts_synthesize(required): The core function accepting text, output path, and voice parameterstts_check(optional): Validation function verifying prerequisites and API connectivitytts_install_help(optional): Function displaying setup instructions when dependencies are missing
This architecture decouples the orchestration logic from vendor-specific implementations, allowing the skill to treat all TTS services as interchangeable black boxes.
Switching Between Providers
Changing TTS providers requires no code changes—only environment variable modifications.
Use the default MiniMax provider without additional configuration:
npm run synthesize-audio
Switch to OpenAI by setting the provider variable:
PRESENTATION_TTS=openai npm run synthesize-audio
Configure ElevenLabs with a specific voice ID:
PRESENTATION_TTS=elevenlabs PRESENTATION_TTS_VOICE=21m00Tcm4TlvDq8ikWAM \
npm run synthesize-audio
For debugging individual providers without the full pipeline, source the script directly and invoke its functions:
source skills/web-video-presentation/templates/scripts/tts-providers/edge-tts.sh
tts_check && tts_synthesize "测试一下" /tmp/test.mp3 "zh-CN-YunxiNeural"
Creating Custom TTS Providers
Developers can extend the skill by creating new provider scripts in skills/web-video-presentation/templates/scripts/tts-providers/. A valid provider requires only a shell script named <provider-name>.sh implementing the tts_synthesize function with this signature:
tts_synthesize() {
local text="$1"
local output_file="$2"
local voice="$3"
# Implementation-specific logic here
}
Optional helper functions tts_check and tts_install_help improve user experience by validating prerequisites and displaying setup guidance. Once the file is placed in the providers directory, set PRESENTATION_TTS=<provider-name> to activate it.
Summary
- The web-video-presentation skill supports MiniMax (default) and OpenAI as production-ready providers, plus ElevenLabs, edge-tts, macOS say, Azure Speech, and Google Cloud TTS as sample implementations.
- The pluggable architecture uses the
PRESENTATION_TTSenvironment variable to dynamically load provider scripts fromskills/web-video-presentation/templates/scripts/tts-providers/. - Each provider must implement
tts_synthesizeand may optionally implementtts_checkandtts_install_help. - Switching providers requires no code changes—only environment variable configuration.
- Custom providers follow a simple three-function shell script contract defined in the README.
Frequently Asked Questions
How do I add a new TTS provider to the web-video-presentation skill?
Create a shell script in skills/web-video-presentation/templates/scripts/tts-providers/ named <your-provider>.sh that implements the tts_synthesize function accepting text, output path, and voice parameters. Optionally add tts_check for validation and tts_install_help for setup instructions. Set PRESENTATION_TTS=your-provider to activate it.
Which TTS provider works without an API key?
The edge-tts provider requires no API key and uses Microsoft's free Edge TTS service. The macOS say provider also requires no cloud credentials but only works on macOS systems and needs ffmpeg for audio format conversion.
Can I use Azure OpenAI instead of the standard OpenAI endpoint?
Yes. The openai.sh provider supports Azure OpenAI endpoints. Configure the OPENAI_API_KEY environment variable and ensure your Azure deployment details are properly set in the provider configuration or through additional environment variables handled by the OpenAI client.
What happens if the TTS provider prerequisites are missing?
If a provider implements the optional tts_check function, the runner script will detect missing dependencies before attempting synthesis. If tts_install_help is implemented, the skill will display platform-specific setup instructions. Without these helpers, the script will fail when attempting to execute the vendor's CLI tool or API client.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →