# Workflow for the Web-Video-Presentation Skill: A Complete 4-Phase Guide

> Master the web-video-presentation skill with a complete 4-phase guide. Learn the workflow from content to final recording. Transform articles into engaging web videos efficiently.

- Repository: [ConardLi/garden-skills](https://github.com/ConardLi/garden-skills)
- Tags: how-to-guide
- Published: 2026-08-30

---

**The web-video-presentation skill transforms written articles into interactive, click-driven 16:9 web videos through a structured four-phase workflow with mandatory checkpoints between content authoring, web development, optional audio synthesis, and final recording.**

The **web-video-presentation** skill in the `ConardLi/garden-skills` repository converts static text into polished, interactive web presentations. According to the source code, this workflow enforces a "double-source principle" that synchronizes script, outline, visuals, and audio through hard checkpoints, ensuring the final deliverable matches the original content exactly.

## Overview of the Four-Phase Workflow

The end-to-end process divides creation into four distinct phases, each terminating in a checkpoint that validates alignment before proceeding:

| Phase | Goal | Key Output |
|-------|------|------------|
| **Phase 1** – Content Authoring | Generate script and outline from source material | [`script.md`](https://github.com/ConardLi/garden-skills/blob/main/script.md), [`outline.md`](https://github.com/ConardLi/garden-skills/blob/main/outline.md), Checkpoint Plan |
| **Phase 2** – Web Development | Scaffold Vite + React project and build chapters | Completed chapters, Checkpoint Audio decision |
| **Phase 3** – Audio Synthesis (Optional) | Produce synchronized voice-over files | MP3 files in `public/audio/` |
| **Phase 4** – Recording & Post-Production | Capture final video via auto or manual mode | Final video file |

## Phase 1 — Content Authoring

This phase establishes the foundation by analyzing input material and generating structured documentation.

### Input Identification and Document Generation

First, the system identifies whether the user provides an article URL, a raw script draft, or just a topic request. Based on this input, it generates two critical files simultaneously:

- **[`script.md`](https://github.com/ConardLi/garden-skills/blob/main/script.md)** – A platform-aware voice script formatted according to [[`references/SCRIPT-STYLE.md`](https://github.com/ConardLi/garden-skills/blob/main/references/SCRIPT-STYLE.md)](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/SCRIPT-STYLE.md)
- **[`outline.md`](https://github.com/ConardLi/garden-skills/blob/main/outline.md)** – A chapter and step plan following [[`references/OUTLINE-FORMAT.md`](https://github.com/ConardLi/garden-skills/blob/main/references/OUTLINE-FORMAT.md)](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/OUTLINE-FORMAT.md)

### Checkpoint Plan

Before advancing, the **Checkpoint Plan** forces validation of five items:

1. The generated **script** accuracy
2. The **outline** structure and flow
3. Three recommended **themes** (derived from `themes/*/theme.json`)
4. A **material checklist** for required assets
5. The chosen **development mode** (A, B, or C)

The user must explicitly confirm or edit each item before the workflow proceeds to scaffolding.

## Phase 2 — Web Development

This phase converts the approved outline into a functioning React application with strict chapter-by-chapter controls.

### Project Scaffolding

Execute the scaffold script to generate a Vite + React project skeleton:

```bash
bash scripts/scaffold.sh ./presentation --theme=warm-keynote

```

This creates the `presentation/` directory with the selected theme's styling tokens, component library, and development toolchain preconfigured.

### Chapter 1 Anchor Implementation

**Chapter 1** serves as the visual and interactive anchor for the entire presentation. The agent must implement this chapter completely—including visuals, narration steps, and assets—before building subsequent content. The user must approve Chapter 1 before any further development begins.

### Development Modes for Remaining Chapters

After Chapter 1 approval, the agent builds remaining chapters using one of three modes defined in the workflow:

- **Mode A – Per-chapter approval (default):** Build each chapter individually, pausing for user validation after every chapter
- **Mode B – Sequential:** Build all remaining chapters in order, then conduct a single comprehensive validation pass
- **Mode C – Parallel (sub-agent):** Dispatch sub-agents to develop multiple chapters concurrently for faster turnaround

### Single Source of Truth

All chapter development references **[`narrations.ts`](https://github.com/ConardLi/garden-skills/blob/main/narrations.ts)** as the sole source of truth for step count and narration text. The global step counter in [`presentation/src/hooks/useStepper.ts`](https://github.com/ConardLi/garden-skills/blob/main/presentation/src/hooks/useStepper.ts) (persisted via `STORAGE_KEY`) updates automatically whenever chapters are added or removed via [`presentation/src/registry/chapters.ts`](https://github.com/ConardLi/garden-skills/blob/main/presentation/src/registry/chapters.ts).

### Checkpoint Audio

Upon completing all chapters, the **Checkpoint Audio** decision point determines whether to proceed with voice synthesis or skip to manual recording.

## Phase 3 — Audio Synthesis (Optional)

This phase generates synchronized voice-overs when automatic narration is required.

### Extracting Narration Segments

Run the extraction script to scan all chapter narration files:

```bash
cd presentation
npm run extract-narrations

```

This command processes every `src/chapters/**/narrations.ts` file and outputs [`audio-segments.json`](https://github.com/ConardLi/garden-skills/blob/main/audio-segments.json) containing timing and text data.

### Synthesizing Audio

Generate MP3 files using the default **minimax** provider:

```bash
npm run synthesize-audio

```

To switch to OpenAI TTS, set the environment variable before running:

```bash
export PRESENTATION_TTS=openai
npm run synthesize-audio

```

Custom TTS providers can be added under `scripts/tts-providers/` following the integration guide in [[`templates/scripts/tts-providers/README.md`](https://github.com/ConardLi/garden-skills/blob/main/templates/scripts/tts-providers/README.md)](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/README.md).

## Phase 4 — Recording & Post-Production

The final phase captures the presentation as a video file using one of two methods.

### Auto-Recording Mode

When audio synthesis is complete, open the auto-play URL:

```bash
open http://localhost:5173/?auto=1

```

Press **Space** to trigger automatic playback of the entire presentation with synchronized audio, then capture the screen in a single take using any screen-recording tool.

### Manual Recording Mode

Without synthesized audio, step through the presentation manually using mouse clicks to advance slides, record the screen capture, and add voice-over in post-production if needed.

No further code changes are required after recording completes.

## Key Files and Architecture

Understanding the file structure ensures proper customization and debugging:

- **[`SKILL.md`](https://github.com/ConardLi/garden-skills/blob/main/SKILL.md)** – Master workflow definition describing phases, checkpoints, and the double-source principle
- **[`scripts/scaffold.sh`](https://github.com/ConardLi/garden-skills/blob/main/scripts/scaffold.sh)** – Bash script generating the Vite + React project with theme selection
- **[`references/CHAPTER-CRAFT.md`](https://github.com/ConardLi/garden-skills/blob/main/references/CHAPTER-CRAFT.md)** – Ten design principles and hard rules for chapter construction
- **[`references/THEMES.md`](https://github.com/ConardLi/garden-skills/blob/main/references/THEMES.md)** – Theme token contract and built-in theme catalog
- **[`presentation/src/registry/chapters.ts`](https://github.com/ConardLi/garden-skills/blob/main/presentation/src/registry/chapters.ts)** – Chapter registration file; edit when adding or removing chapters
- **[`presentation/src/hooks/useStepper.ts`](https://github.com/ConardLi/garden-skills/blob/main/presentation/src/hooks/useStepper.ts)** – Global step counter hook managing presentation state
- **[`references/RECORDING.md`](https://github.com/ConardLi/garden-skills/blob/main/references/RECORDING.md)** – Detailed instructions for auto and manual screen recording workflows

## Summary

- The **web-video-presentation** skill uses a rigid four-phase workflow to convert articles into interactive 16:9 web videos.
- **Checkpoint Plan** and **Checkpoint Audio** enforce mandatory validation gates between phases.
- Development occurs in **Mode A (per-chapter)**, **Mode B (sequential)**, or **Mode C (parallel)** after establishing a Chapter 1 anchor.
- **[`narrations.ts`](https://github.com/ConardLi/garden-skills/blob/main/narrations.ts)** serves as the single source of truth for both visual steps and audio synthesis.
- Audio synthesis supports **minimax** (default) and **OpenAI TTS** providers via environment configuration.
- Final delivery uses either **auto-recording** (`?auto=1`) with synthesized audio or manual screen capture.

## Frequently Asked Questions

### What is the "double-source principle" in the web-video-presentation workflow?

The double-source principle ensures that the **script**, **outline**, **visual design**, and **audio** remain perfectly synchronized throughout development. By enforcing checkpoints between phases and using [`narrations.ts`](https://github.com/ConardLi/garden-skills/blob/main/narrations.ts) as the single source of truth for step count and text, the workflow prevents drift between the original content and the final presentation.

### How do I switch from the default minimax TTS provider to OpenAI's text-to-speech?

Set the `PRESENTATION_TTS` environment variable to `openai` before running the synthesis command. Ensure your `OPENAI_API_KEY` is configured in your environment, then execute `npm run synthesize-audio`. The system will route audio generation through OpenAI's API instead of the default minimax provider.

### What happens if I need to add or remove chapters after starting development?

Edit [`presentation/src/registry/chapters.ts`](https://github.com/ConardLi/garden-skills/blob/main/presentation/src/registry/chapters.ts) to register new chapters or remove existing ones. If removing the demo chapter generated by scaffolding, delete both the chapter directory (`rm -rf presentation/src/chapters/01-example`) and its corresponding entry in the registry file. The `useStepper` hook automatically recalculates the global step count based on the updated registry.

### Can I record the presentation without synthesizing audio first?

Yes. While auto-recording mode (`?auto=1`) requires synthesized audio to function, you can use **manual recording mode** instead. Navigate to the presentation URL, manually click through each step at your preferred pace, and capture the screen with any recording software. You can add voice-over during post-production or leave it as a silent presentation.