Workflow for the Web-Video-Presentation Skill: A Complete 4-Phase Guide
The web-video-presentation skill transforms written articles into interactive, click-driven 16:9 web videos through a structured four-phase workflow with mandatory checkpoints between content authoring, web development, optional audio synthesis, and final recording.
The web-video-presentation skill in the ConardLi/garden-skills repository converts static text into polished, interactive web presentations. According to the source code, this workflow enforces a "double-source principle" that synchronizes script, outline, visuals, and audio through hard checkpoints, ensuring the final deliverable matches the original content exactly.
Overview of the Four-Phase Workflow
The end-to-end process divides creation into four distinct phases, each terminating in a checkpoint that validates alignment before proceeding:
| Phase | Goal | Key Output |
|---|---|---|
| Phase 1 – Content Authoring | Generate script and outline from source material | script.md, outline.md, Checkpoint Plan |
| Phase 2 – Web Development | Scaffold Vite + React project and build chapters | Completed chapters, Checkpoint Audio decision |
| Phase 3 – Audio Synthesis (Optional) | Produce synchronized voice-over files | MP3 files in public/audio/ |
| Phase 4 – Recording & Post-Production | Capture final video via auto or manual mode | Final video file |
Phase 1 — Content Authoring
This phase establishes the foundation by analyzing input material and generating structured documentation.
Input Identification and Document Generation
First, the system identifies whether the user provides an article URL, a raw script draft, or just a topic request. Based on this input, it generates two critical files simultaneously:
script.md– A platform-aware voice script formatted according to [references/SCRIPT-STYLE.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/SCRIPT-STYLE.md)outline.md– A chapter and step plan following [references/OUTLINE-FORMAT.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/OUTLINE-FORMAT.md)
Checkpoint Plan
Before advancing, the Checkpoint Plan forces validation of five items:
- The generated script accuracy
- The outline structure and flow
- Three recommended themes (derived from
themes/*/theme.json) - A material checklist for required assets
- The chosen development mode (A, B, or C)
The user must explicitly confirm or edit each item before the workflow proceeds to scaffolding.
Phase 2 — Web Development
This phase converts the approved outline into a functioning React application with strict chapter-by-chapter controls.
Project Scaffolding
Execute the scaffold script to generate a Vite + React project skeleton:
bash scripts/scaffold.sh ./presentation --theme=warm-keynote
This creates the presentation/ directory with the selected theme's styling tokens, component library, and development toolchain preconfigured.
Chapter 1 Anchor Implementation
Chapter 1 serves as the visual and interactive anchor for the entire presentation. The agent must implement this chapter completely—including visuals, narration steps, and assets—before building subsequent content. The user must approve Chapter 1 before any further development begins.
Development Modes for Remaining Chapters
After Chapter 1 approval, the agent builds remaining chapters using one of three modes defined in the workflow:
- Mode A – Per-chapter approval (default): Build each chapter individually, pausing for user validation after every chapter
- Mode B – Sequential: Build all remaining chapters in order, then conduct a single comprehensive validation pass
- Mode C – Parallel (sub-agent): Dispatch sub-agents to develop multiple chapters concurrently for faster turnaround
Single Source of Truth
All chapter development references narrations.ts as the sole source of truth for step count and narration text. The global step counter in presentation/src/hooks/useStepper.ts (persisted via STORAGE_KEY) updates automatically whenever chapters are added or removed via presentation/src/registry/chapters.ts.
Checkpoint Audio
Upon completing all chapters, the Checkpoint Audio decision point determines whether to proceed with voice synthesis or skip to manual recording.
Phase 3 — Audio Synthesis (Optional)
This phase generates synchronized voice-overs when automatic narration is required.
Extracting Narration Segments
Run the extraction script to scan all chapter narration files:
cd presentation
npm run extract-narrations
This command processes every src/chapters/**/narrations.ts file and outputs audio-segments.json containing timing and text data.
Synthesizing Audio
Generate MP3 files using the default minimax provider:
npm run synthesize-audio
To switch to OpenAI TTS, set the environment variable before running:
export PRESENTATION_TTS=openai
npm run synthesize-audio
Custom TTS providers can be added under scripts/tts-providers/ following the integration guide in [templates/scripts/tts-providers/README.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/README.md).
Phase 4 — Recording & Post-Production
The final phase captures the presentation as a video file using one of two methods.
Auto-Recording Mode
When audio synthesis is complete, open the auto-play URL:
open http://localhost:5173/?auto=1
Press Space to trigger automatic playback of the entire presentation with synchronized audio, then capture the screen in a single take using any screen-recording tool.
Manual Recording Mode
Without synthesized audio, step through the presentation manually using mouse clicks to advance slides, record the screen capture, and add voice-over in post-production if needed.
No further code changes are required after recording completes.
Key Files and Architecture
Understanding the file structure ensures proper customization and debugging:
SKILL.md– Master workflow definition describing phases, checkpoints, and the double-source principlescripts/scaffold.sh– Bash script generating the Vite + React project with theme selectionreferences/CHAPTER-CRAFT.md– Ten design principles and hard rules for chapter constructionreferences/THEMES.md– Theme token contract and built-in theme catalogpresentation/src/registry/chapters.ts– Chapter registration file; edit when adding or removing chapterspresentation/src/hooks/useStepper.ts– Global step counter hook managing presentation statereferences/RECORDING.md– Detailed instructions for auto and manual screen recording workflows
Summary
- The web-video-presentation skill uses a rigid four-phase workflow to convert articles into interactive 16:9 web videos.
- Checkpoint Plan and Checkpoint Audio enforce mandatory validation gates between phases.
- Development occurs in Mode A (per-chapter), Mode B (sequential), or Mode C (parallel) after establishing a Chapter 1 anchor.
narrations.tsserves as the single source of truth for both visual steps and audio synthesis.- Audio synthesis supports minimax (default) and OpenAI TTS providers via environment configuration.
- Final delivery uses either auto-recording (
?auto=1) with synthesized audio or manual screen capture.
Frequently Asked Questions
What is the "double-source principle" in the web-video-presentation workflow?
The double-source principle ensures that the script, outline, visual design, and audio remain perfectly synchronized throughout development. By enforcing checkpoints between phases and using narrations.ts as the single source of truth for step count and text, the workflow prevents drift between the original content and the final presentation.
How do I switch from the default minimax TTS provider to OpenAI's text-to-speech?
Set the PRESENTATION_TTS environment variable to openai before running the synthesis command. Ensure your OPENAI_API_KEY is configured in your environment, then execute npm run synthesize-audio. The system will route audio generation through OpenAI's API instead of the default minimax provider.
What happens if I need to add or remove chapters after starting development?
Edit presentation/src/registry/chapters.ts to register new chapters or remove existing ones. If removing the demo chapter generated by scaffolding, delete both the chapter directory (rm -rf presentation/src/chapters/01-example) and its corresponding entry in the registry file. The useStepper hook automatically recalculates the global step count based on the updated registry.
Can I record the presentation without synthesizing audio first?
Yes. While auto-recording mode (?auto=1) requires synthesized audio to function, you can use manual recording mode instead. Navigate to the presentation URL, manually click through each step at your preferred pace, and capture the screen with any recording software. You can add voice-over during post-production or leave it as a silent presentation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →