Built-in TTS Providers for Garden Skills: MiniMax and OpenAI Guide
Garden Skills ships with two built-in Text-to-Speech providers—MiniMax and OpenAI—that implement a standard shell interface for the Web Video Presentation skill.
Garden Skills (ConardLi/garden-skills) is an open-source automation framework for generating video presentations from web content. Its built-in TTS providers deliver plug-and-play speech synthesis through a modular shell-script architecture, allowing developers to toggle between Chinese-optimized and multilingual voice engines using environment variables.
Architecture of Garden Skills TTS Providers
Every TTS provider in Garden Skills follows a strict contract defined in the Web Video Presentation skill. Each provider is a shell script located in skills/web-video-presentation/templates/scripts/tts-providers/ that exposes three standard functions:
tts_check: Validates prerequisites such as CLI binaries or API keystts_install_help: (Optional) Returns installation instructions iftts_checkfailstts_synthesize: (Required) Executes the actual text-to-speech conversion
The adapter script skills/web-video-presentation/templates/scripts/synthesize-audio.sh loads the active provider via the PRESENTATION_TTS environment variable, defaulting to minimax if unspecified:
PRESENTATION_TTS=${PRESENTATION_TTS:-minimax}
source "$SCRIPT_DIR/tts-providers/$PROVIDER.sh"
If the provider's tts_check function reports a failure, the runner prints the helper text from tts_install_help and aborts execution.
MiniMax TTS Provider
The MiniMax provider is the default engine, optimized for high-quality Chinese narration.
Implementation resides in skills/web-video-presentation/templates/scripts/tts-providers/minimax.sh. It invokes the mmx CLI (mmx-cli) to generate speech in a single call, offering extensive voice options tailored for Mandarin content.
To synthesize audio with MiniMax:
npm run synthesize-audio -- --text "你好,欢迎观看。"
OpenAI TTS Provider
The OpenAI provider connects to the OpenAI Audio Speech REST API via curl, supporting multiple voices and HD model variants.
Located at skills/web-video-presentation/templates/scripts/tts-providers/openai.sh, this provider accepts the OPENAI_TTS_MODEL environment variable (e.g., tts-1 or tts-1-hd) and the PRESENTATION_TTS_VOICE variable to select from: alloy, echo, fable, onyx, nova, and shimmer.
Example with voice selection:
PRESENTATION_TTS=openai PRESENTATION_TTS_VOICE=nova \
npm run synthesize-audio -- --text "Hello, welcome to the presentation."
Configuring and Extending TTS Providers
Toggle between engines by setting PRESENTATION_TTS before invoking the synthesis script. The system validates the provider via tts_check; if validation fails, it displays the helper text from tts_install_help and aborts.
To add custom providers, create a new <name>.sh file in skills/web-video-presentation/templates/scripts/tts-providers/ and implement the required interface functions. The README.md in that directory provides contract specifications and boilerplate snippets for additional services like ElevenLabs.
Example invocation of a custom provider:
PRESENTATION_TTS=elevenlabs npm run synthesize-audio -- --text "Custom TTS test."
Summary
- Garden Skills provides two built-in TTS providers: MiniMax (default, Chinese-optimized) and OpenAI (multilingual, six voices).
- Providers are shell scripts in
skills/web-video-presentation/templates/scripts/tts-providers/implementingtts_checkandtts_synthesize. - Selection occurs via the
PRESENTATION_TTSenvironment variable parsed bysynthesize-audio.sh. - OpenAI supports
PRESENTATION_TTS_VOICEandOPENAI_TTS_MODELfor fine-tuning output quality. - The architecture supports custom providers through a standardized three-function interface.
Frequently Asked Questions
How do I switch between TTS providers in Garden Skills?
Set the PRESENTATION_TTS environment variable to either minimax or openai before running npm run synthesize-audio. If undefined, the system defaults to MiniMax as implemented in skills/web-video-presentation/templates/scripts/synthesize-audio.sh.
What voices are available with the OpenAI TTS provider?
The OpenAI provider supports six voices: alloy, echo, fable, onyx, nova, and shimmer. Specify your choice via the PRESENTATION_TTS_VOICE environment variable.
Do I need to install anything extra to use the MiniMax provider?
Yes. The MiniMax provider requires the mmx CLI (mmx-cli) to be installed and available in your PATH. The tts_check function in minimax.sh validates this dependency before synthesis begins.
Can I add my own TTS provider to Garden Skills?
Yes. Create a new shell script in skills/web-video-presentation/templates/scripts/tts-providers/ implementing tts_check and tts_synthesize (optionally tts_install_help), then invoke it by setting PRESENTATION_TTS to your script's basename.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →