How to Synthesize Audio Using Custom TTS Providers for the Web-Video-Presentation Skill

You synthesize audio in the web-video-presentation skill by implementing a shell script that conforms to the tts_synthesize contract, placing it in scripts/tts-providers/, and running synthesize-audio.sh --provider <name> to generate MP3 files from your narration text.

The web-video-presentation skill in the ConardLi/garden-skills repository generates automated slideshow videos where each step can include synthesized narration. The architecture separates content definition from audio generation through a pipeline that extracts text, delegates synthesis to pluggable providers, and serves static MP3 assets at runtime.

Audio Synthesis Architecture

The synthesis workflow follows a four-stage pipeline defined in the skill templates.

Narration Sources

Each chapter defines its spoken content in a narrations.ts file that exports a flat array of strings. Each array index corresponds to a presentation step; empty strings indicate silent steps that skip synthesis. In skills/web-video-presentation/templates/src/chapters/01-example/narrations.ts, the default template exports a typed array:

export default [
  "Welcome to this demonstration",
  "", // Silent step
  "Next we will explore the concept",
] as const;

Extraction Pipeline

The extract-narrations.ts script scans the chapter registry, dynamically imports every narrations.ts module, filters out empty strings, and serializes a flat list of segments to audio-segments.json. Located at skills/web-video-presentation/templates/scripts/extract-narrations.ts, this utility maps each non-empty narration to a unique identifier combining the chapter name and step index.

Synthesis Driver

The synthesize-audio.sh shell script drives the actual TTS generation. It reads audio-segments.json, iterates over segments, and delegates synthesis to a provider script selected via the --provider flag. The driver manages output directories and writes final MP3 files to public/audio/<chapter>/<step>.mp3. This script is located at skills/web-video-presentation/templates/scripts/synthesize-audio.sh.

Runtime Playback

At runtime, the useAudioPlayer hook in skills/web-video-presentation/templates/src/hooks/useAudioPlayer.ts creates a hidden HTMLAudioElement, constructs the URL for the current step’s MP3 file, and manages playback state. If the MP3 is missing—such as for silent steps or failed synthesis—the hook falls back to a timer-based estimate to maintain presentation flow.

Implementing a Custom TTS Provider

Adding support for Azure, Google Cloud, or a private TTS API requires implementing a contract-defined shell interface.

The Provider Contract

The provider contract is documented in skills/web-video-presentation/templates/scripts/tts-providers/README.md. Every provider script must export a tts_synthesize function accepting two positional arguments:

tts_synthesize() {
  local text="$1"      # The narration text to synthesize

  local output_path="$2" # Absolute path where the MP3 must be written

  # Implementation must write a valid MP3 file to $output_path

}

Reference implementations exist for OpenAI (openai.sh) and Minimax (minimax.sh) in the same directory, demonstrating HTTP request patterns and error handling.

Step-by-Step Implementation

  1. Create the provider file inside skills/web-video-presentation/templates/scripts/tts-providers/, naming it <provider-name>.sh.

  2. Implement the synthesis function following the contract signature. Wrap API calls, handle retries, and ensure the output file is a valid MP3 encoded at a web-compatible bitrate.

  3. Set executable permissions on the script:

    chmod +x scripts/tts-providers/my-custom.sh
  4. Register the provider by passing the basename to the synthesis driver:

    ./scripts/synthesize-audio.sh --provider my-custom

Execution and Verification

Run the extraction script first to ensure audio-segments.json is current:

cd skills/web-video-presentation/templates
node scripts/extract-narrations.js

Execute synthesis:

./scripts/synthesize-audio.sh --provider my-custom

Verify output by checking for MP3 files in public/audio/<chapter>/. The driver skips existing files on subsequent runs, allowing incremental updates.

Practical Implementation Examples

Extract narrations before synthesizing:

cd skills/web-video-presentation/templates
node scripts/extract-narrations.js

Synthesize using the built-in Minimax provider:

./scripts/synthesize-audio.sh --provider minimax

Synthesize using a custom Azure provider:

./scripts/synthesize-audio.sh --provider azure-tts

Accessing audio in a chapter component:

import { useAudioPlayer } from "../hooks/useAudioPlayer";

export default function ChapterStep({ step, chapter }: ChapterStepProps) {
  const { play, isPlaying, error } = useAudioPlayer(step, chapter);
  
  return (
    <button onClick={play} disabled={isPlaying}>
      {isPlaying ? "Playing..." : "Play Narration"}
    </button>
  );
}

Summary

  • Content definition occurs in chapter-level narrations.ts files that export string arrays.
  • Extraction happens via extract-narrations.ts, which produces audio-segments.json.
  • Synthesis is driven by synthesize-audio.sh, which delegates to pluggable provider scripts implementing the tts_synthesize contract.
  • Custom providers reside in scripts/tts-providers/ and must write MP3 files to the path provided as the second argument.
  • Playback is handled automatically by the useAudioPlayer hook, which falls back to timed advances if audio files are absent.

Frequently Asked Questions

What is the exact function signature my custom TTS provider must implement?

Your provider script must define a tts_synthesize function that accepts exactly two parameters: the text string to synthesize and the absolute file system path where the MP3 must be saved. The function should not return a value; success is determined by the existence of a valid MP3 file at the specified path after execution.

Where do I place my custom provider script in the repository?

Place the shell script inside skills/web-video-presentation/templates/scripts/tts-providers/ and ensure the filename matches the identifier you pass to the --provider flag (e.g., custom.sh invoked with --provider custom). The script must have executable permissions (chmod +x) to be invoked by the synthesis driver.

How does the presentation handle missing or failed audio synthesis?

The useAudioPlayer hook checks for the existence of the MP3 file at public/audio/<chapter>/<step>.mp3. If the file is missing—either because the narration was an empty string or synthesis failed—the hook automatically falls back to a duration estimate based on text length, ensuring the presentation advances without breaking the user experience.

Can I use different TTS providers for different chapters?

The current architecture processes all chapters in a single invocation using one provider. To use different providers per chapter, you would need to run synthesize-audio.sh separately for each chapter subdirectory with different --provider values, or modify the extraction script to include provider metadata in audio-segments.json and update the driver to route segments conditionally.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →