How the Web-Video-Presentation Skill Converts Articles into Interactive Presentations

The web-video-presentation skill transforms markdown articles into 16:9 interactive web applications through a four-phase pipeline that generates voice-over scripts, scaffolds a Vite+React codebase, and produces click-driven presentations with optional synthesized audio.

The web-video-presentation skill in the ConardLi/garden-skills repository provides a declarative, checkpoint-driven workflow for converting static content into dynamic visual narratives. This open-source tool treats each article as a blueprint for a self-contained web application that behaves like a video, advancing step-by-step through user interactions or automated playback. The entire process is iterative, with hard checkpoints that ensure alignment between the original content and the final presentation before any code is generated.

Phase 1: Content Generation

The pipeline begins by analyzing the source material—either a raw article.md or an existing script—and producing two critical artifacts in a single reasoning pass.

These documents serve as the single source of truth for timing, with each step representing one beat of the voice-over that drives the entire presentation flow.

Checkpoint Validation Before Code Generation

Before generating any code, the skill enforces a checkpoint plan that requires user confirmation of five elements:

  1. The generated script
  2. The development outline
  3. The visual theme selection
  4. Required asset inventory
  5. Development mode preferences

This validation ensures downstream scaffolding operates on a stable foundation, preventing costly refactoring later in the pipeline.

Phase 2: Web Scaffold and Chapter Development

Project Initialization with scaffold.sh

The scripts/scaffold.sh script generates a fresh Vite, React, and TypeScript project under the presentation/ directory using the selected theme via the --theme=<id> flag. The scaffold includes a demo chapter (01-example) that developers must remove before implementing real content.


# Scaffold a new presentation with the modern theme

bash /path/to/web-video-presentation/scripts/scaffold.sh ./presentation --theme=modern

# Remove the demo chapter and its registry entry

rm -rf presentation/src/chapters/01-example

# Edit presentation/src/registry/chapters.ts to remove EXAMPLE_CHAPTER

Chapter Architecture and Step Logic

Each chapter resides in presentation/src/chapters/<NN>-<id>/ and contains three files:

  • A .tsx view component
  • A .css stylesheet
  • A narrations.ts file (the single source of truth for step count and voice text)

The global step counter (useStepper in templates/src/hooks/useStepper.ts) drives the UI deterministically. The rendering logic follows a simple conditional pattern:

if (step === N) return <FullScene />;

Users advance through steps via click interactions, creating a controlled, pace-driven experience. The implementation adheres to the CHAPTER-CRAFT principles outlined in [CHAPTER-CRAFT.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/CHAPTER-CRAFT.md), including:

  • The "one-item = one-step" reveal rule
  • The "double-source" principle (synchronizing script, article, and visual elements)
  • Ten design constraints, including fixed 16:9 stage ratio, hidden controls, and content-driven animations

Phase 3: Audio Synthesis Pipeline

The skill includes a provider-agnostic text-to-speech (TTS) system for generating voice-overs. The process involves two npm scripts:

By default, the system uses MiniMax (mmx-cli), but developers can switch to OpenAI TTS by setting an environment variable:


# Default MiniMax synthesis

npm run synthesize-audio

# Alternative OpenAI TTS provider

PRESENTATION_TTS=openai npm run synthesize-audio

Synthesized MP3 files are stored under public/audio/<id>/<N>.mp3 and played via the useAudioPlayer hook. Additional providers are documented in templates/scripts/tts-providers/README.md.

Phase 4: Recording and Post-Production

The generated web application supports two capture modes for final video production:

  • Auto mode: Accessing localhost:5173/?auto=1 automatically advances through the entire step sequence, enabling one-click screen capture for smooth, continuous video output.
  • Manual mode: Users click through steps at their preferred pace, allowing for natural pauses and emphasis that can be refined in post-production with any video editor.

Core Architecture and Key Files

The web-video-presentation skill relies on a specific file structure and hook architecture:

File Function
[SKILL.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/SKILL.md) High-level workflow overview and hard rules
templates/src/hooks/useStepper.ts Global step counter with persistent storage
templates/src/hooks/useAudioPlayer.ts Per-step audio playback management
templates/src/registry/chapters.ts Registry wiring chapter modules into the application
templates/scripts/extract-narrations.ts CLI tool building audio-segments.json
scripts/scaffold.sh One-click project scaffolder

To run the development server and preview chapters:

cd presentation
npm run dev  # → http://localhost:5173

Summary

  • The web-video-presentation skill converts article.md into interactive 16:9 web applications through a four-phase pipeline.
  • Content Generation produces script.md and outline.md as the single source of truth for timing and structure.
  • Checkpoint validation requires user confirmation of five elements before code generation begins.
  • Chapter Development uses scaffold.sh to create Vite+React projects where narrations.ts controls step logic via useStepper.
  • Audio Synthesis supports MiniMax (default) and OpenAI TTS through configurable npm scripts.
  • Recording modes include automatic playback (?auto=1) for hands-free capture or manual clicking for controlled pacing.

Frequently Asked Questions

What input format does the web-video-presentation skill require?

The skill accepts a markdown article file (article.md) or an existing voice-over script. During Phase 1, it processes this input to generate script.md and outline.md according to the style contracts in SCRIPT-STYLE.md and OUTLINE-FORMAT.md, ensuring the content follows a specific narrative structure before development begins.

How does the stepper system control presentation flow?

The useStepper hook (located in templates/src/hooks/useStepper.ts) maintains a global step counter that determines which scene renders based on the current step number. Each chapter's narrations.ts file defines the total step count and corresponding voice text, creating a deterministic sequence where if (step === N) return <FullScene/> logic governs visual reveals.

Can I use different text-to-speech providers?

Yes. While the default configuration uses MiniMax (mmx-cli), you can switch to OpenAI TTS by setting the PRESENTATION_TTS=openai environment variable before running npm run synthesize-audio. The architecture in templates/scripts/tts-providers/ supports additional provider extensions through a standardized interface.

What is the "one-item = one-step" principle?

This principle, defined in CHAPTER-CRAFT.md, mandates that each visual element or content item corresponds to exactly one step in the narration sequence. This rule ensures tight synchronization between the voice-over and visual reveals, preventing cognitive overload and maintaining the presentation's rhythmic pacing as users click through the stepper.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →