How the Web-Video-Presentation Skill Converts Articles into Interactive Presentations
The web-video-presentation skill transforms markdown articles into 16:9 interactive web applications through a four-phase pipeline that generates voice-over scripts, scaffolds a Vite+React codebase, and produces click-driven presentations with optional synthesized audio.
The web-video-presentation skill in the ConardLi/garden-skills repository provides a declarative, checkpoint-driven workflow for converting static content into dynamic visual narratives. This open-source tool treats each article as a blueprint for a self-contained web application that behaves like a video, advancing step-by-step through user interactions or automated playback. The entire process is iterative, with hard checkpoints that ensure alignment between the original content and the final presentation before any code is generated.
Phase 1: Content Generation
The pipeline begins by analyzing the source material—either a raw article.md or an existing script—and producing two critical artifacts in a single reasoning pass.
script.md: A platform-agnostic voice-over script that preserves the original language while following strict style contracts defined in [SCRIPT-STYLE.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/SCRIPT-STYLE.md).outline.md: A development plan that segments content into chapters, steps, and information pools, formatted according to [OUTLINE-FORMAT.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/OUTLINE-FORMAT.md).
These documents serve as the single source of truth for timing, with each step representing one beat of the voice-over that drives the entire presentation flow.
Checkpoint Validation Before Code Generation
Before generating any code, the skill enforces a checkpoint plan that requires user confirmation of five elements:
- The generated script
- The development outline
- The visual theme selection
- Required asset inventory
- Development mode preferences
This validation ensures downstream scaffolding operates on a stable foundation, preventing costly refactoring later in the pipeline.
Phase 2: Web Scaffold and Chapter Development
Project Initialization with scaffold.sh
The scripts/scaffold.sh script generates a fresh Vite, React, and TypeScript project under the presentation/ directory using the selected theme via the --theme=<id> flag. The scaffold includes a demo chapter (01-example) that developers must remove before implementing real content.
# Scaffold a new presentation with the modern theme
bash /path/to/web-video-presentation/scripts/scaffold.sh ./presentation --theme=modern
# Remove the demo chapter and its registry entry
rm -rf presentation/src/chapters/01-example
# Edit presentation/src/registry/chapters.ts to remove EXAMPLE_CHAPTER
Chapter Architecture and Step Logic
Each chapter resides in presentation/src/chapters/<NN>-<id>/ and contains three files:
- A
.tsxview component - A
.cssstylesheet - A
narrations.tsfile (the single source of truth for step count and voice text)
The global step counter (useStepper in templates/src/hooks/useStepper.ts) drives the UI deterministically. The rendering logic follows a simple conditional pattern:
if (step === N) return <FullScene />;
Users advance through steps via click interactions, creating a controlled, pace-driven experience. The implementation adheres to the CHAPTER-CRAFT principles outlined in [CHAPTER-CRAFT.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/CHAPTER-CRAFT.md), including:
- The "one-item = one-step" reveal rule
- The "double-source" principle (synchronizing script, article, and visual elements)
- Ten design constraints, including fixed 16:9 stage ratio, hidden controls, and content-driven animations
Phase 3: Audio Synthesis Pipeline
The skill includes a provider-agnostic text-to-speech (TTS) system for generating voice-overs. The process involves two npm scripts:
npm run extract-narrations: Scans allnarrations.tsfiles and generatesaudio-segments.jsonnpm run synthesize-audio: Executes the TTS pipeline
By default, the system uses MiniMax (mmx-cli), but developers can switch to OpenAI TTS by setting an environment variable:
# Default MiniMax synthesis
npm run synthesize-audio
# Alternative OpenAI TTS provider
PRESENTATION_TTS=openai npm run synthesize-audio
Synthesized MP3 files are stored under public/audio/<id>/<N>.mp3 and played via the useAudioPlayer hook. Additional providers are documented in templates/scripts/tts-providers/README.md.
Phase 4: Recording and Post-Production
The generated web application supports two capture modes for final video production:
- Auto mode: Accessing
localhost:5173/?auto=1automatically advances through the entire step sequence, enabling one-click screen capture for smooth, continuous video output. - Manual mode: Users click through steps at their preferred pace, allowing for natural pauses and emphasis that can be refined in post-production with any video editor.
Core Architecture and Key Files
The web-video-presentation skill relies on a specific file structure and hook architecture:
| File | Function |
|---|---|
[SKILL.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/SKILL.md) |
High-level workflow overview and hard rules |
templates/src/hooks/useStepper.ts |
Global step counter with persistent storage |
templates/src/hooks/useAudioPlayer.ts |
Per-step audio playback management |
templates/src/registry/chapters.ts |
Registry wiring chapter modules into the application |
templates/scripts/extract-narrations.ts |
CLI tool building audio-segments.json |
scripts/scaffold.sh |
One-click project scaffolder |
To run the development server and preview chapters:
cd presentation
npm run dev # → http://localhost:5173
Summary
- The web-video-presentation skill converts
article.mdinto interactive 16:9 web applications through a four-phase pipeline. - Content Generation produces
script.mdandoutline.mdas the single source of truth for timing and structure. - Checkpoint validation requires user confirmation of five elements before code generation begins.
- Chapter Development uses
scaffold.shto create Vite+React projects wherenarrations.tscontrols step logic viauseStepper. - Audio Synthesis supports MiniMax (default) and OpenAI TTS through configurable npm scripts.
- Recording modes include automatic playback (
?auto=1) for hands-free capture or manual clicking for controlled pacing.
Frequently Asked Questions
What input format does the web-video-presentation skill require?
The skill accepts a markdown article file (article.md) or an existing voice-over script. During Phase 1, it processes this input to generate script.md and outline.md according to the style contracts in SCRIPT-STYLE.md and OUTLINE-FORMAT.md, ensuring the content follows a specific narrative structure before development begins.
How does the stepper system control presentation flow?
The useStepper hook (located in templates/src/hooks/useStepper.ts) maintains a global step counter that determines which scene renders based on the current step number. Each chapter's narrations.ts file defines the total step count and corresponding voice text, creating a deterministic sequence where if (step === N) return <FullScene/> logic governs visual reveals.
Can I use different text-to-speech providers?
Yes. While the default configuration uses MiniMax (mmx-cli), you can switch to OpenAI TTS by setting the PRESENTATION_TTS=openai environment variable before running npm run synthesize-audio. The architecture in templates/scripts/tts-providers/ supports additional provider extensions through a standardized interface.
What is the "one-item = one-step" principle?
This principle, defined in CHAPTER-CRAFT.md, mandates that each visual element or content item corresponds to exactly one step in the narration sequence. This rule ensures tight synchronization between the voice-over and visual reveals, preventing cognitive overload and maintaining the presentation's rhythmic pacing as users click through the stepper.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →