Web-Video-Presentation Skill Workflow: A 4-Phase Guide to Interactive Web Videos

The web-video-presentation skill workflow converts written articles into interactive, click-driven 16:9 web presentations through four distinct phases—Content Authoring, Web Development, optional Audio Synthesis, and Recording—each guarded by hard checkpoints that enforce alignment between script, visuals, and audio.

The web-video-presentation skill in the ConardLi/garden-skills repository provides a deterministic pipeline for transforming static content into dynamic browser-based videos. By following the workflow defined in [SKILL.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/SKILL.md), users scaffold a Vite + React project, implement chapter-based navigation, and optionally synthesize synchronized voice-overs to produce a final screen-capture-ready artifact.

Phase 1 – Content Authoring

The workflow begins with structured content preparation. The agent identifies whether the user provides an existing article, a raw script, or only a topic seed.

Generate base documents. In a single pass, the system produces two foundational files:

Checkpoint Plan. Before proceeding, the workflow pauses for user validation of five items: the script, the outline, three recommended themes (derived from themes/*/theme.json), the asset checklist, and the development mode selection (A, B, or C). This checkpoint prevents downstream re-work by locking the creative direction early.

Phase 2 – Web Development

With the plan approved, the system scaffolds the presentation engine and builds out each chapter.

Scaffold the project. Run the bootstrap script to create a Vite + React skeleton:

bash scripts/scaffold.sh ./presentation --theme=warm-keynote

This generates the presentation/ directory with the selected theme assets. Remove the demo chapter before starting custom work:

rm -rf presentation/src/chapters/01-example

# Remove EXAMPLE_CHAPTER from presentation/src/registry/chapters.ts

Anchor chapter implementation. Chapter 1 is built first as the anchor—a full-screen proof-of-concept that includes visuals, narration steps, and assets. The user must approve this chapter before the agent continues.

Build remaining chapters. The workflow supports three execution modes defined at Checkpoint Plan:

  • Mode A (Per-chapter approval) – Default. Each chapter is implemented individually and requires explicit user validation before proceeding.
  • Mode B (Sequential) – All remaining chapters are built in order, followed by a single bulk validation pass.
  • Mode C (Parallel/Sub-agent) – Distributes chapter construction across sub-agents for concurrent development.

Single source of truth. Every chapter references narrations.ts (located at presentation/src/chapters/**/narrations.ts) as the authoritative definition of step count and narration text. The global step counter in presentation/src/hooks/useStepper.ts (persisted via STORAGE_KEY) automatically adjusts when chapters are added or removed, ensuring the navigation state remains consistent.

Start the dev server to preview:

cd presentation
npm install
npm run dev   # Serves at http://localhost:5173

Phase 3 – Audio Synthesis (Optional)

Audio generation is deferred until after visual development is complete, respecting the "double-source principle" that keeps script and visuals synchronized.

Extract narration segments. Scan all narrations.ts files and compile timing data:

npm run extract-narrations   # Creates audio-segments.json

Synthesize voice-overs. By default, the system uses the minimax provider (mmx-cli) to generate MP3s:

npm run synthesize-audio     # Generates MP3s under public/audio/

Switch to OpenAI TTS by setting an environment variable:

export PRESENTATION_TTS=openai
npm run synthesize-audio

Custom providers can be added under scripts/tts-providers/ following the integration guide in [templates/scripts/tts-providers/README.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/templates/scripts/tts-providers/README.md).

Checkpoint Audio. After synthesis, the user reviews the generated MP3s. Completion of this checkpoint signals readiness for final recording.

Phase 4 – Recording & Post-Production

The final phase captures the presentation as a video file.

Auto-recording mode. With synthesized audio in place, open the auto-play URL and initiate seamless capture:

open http://localhost:5173/?auto=1

# Press SPACE to auto-play the entire presentation, then capture your screen

The ?auto=1 query parameter triggers sequential step advancement without manual clicking, suitable for one-take screen recording.

Manual recording mode. Step through the presentation manually using click or keyboard navigation, record with any screen-capture tool, and optionally composite the synthesized audio in post-production.

No further code changes are required after recording. The workflow delivers a final video file that reflects the approved script, outline, and visual design.

Key Files and Architecture

Understanding the repository structure ensures correct customization and debugging:

Summary

  • The web-video-presentation skill workflow divides production into four phases: Content Authoring, Web Development, Audio Synthesis, and Recording.
  • Checkpoint Plan and Checkpoint Audio enforce hard stops that validate script, outline, themes, and audio before proceeding.
  • narrations.ts serves as the single source of truth for step count and narration text, while useStepper maintains global navigation state.
  • Development Mode A (per-chapter approval), Mode B (sequential), and Mode C (parallel) offer flexible build strategies.
  • Audio synthesis defaults to Minimax but supports OpenAI TTS via the PRESENTATION_TTS environment variable.
  • Final capture uses either auto-recording (?auto=1) for hands-free playback or manual stepping for granular control.

Frequently Asked Questions

What input formats does the web-video-presentation skill accept?

The workflow accepts three input types: a full article URL or markdown file, a partial draft script, or a simple topic request. In all cases, the system generates script.md and outline.md following the style guides in [references/SCRIPT-STYLE.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/SCRIPT-STYLE.md) and [references/OUTLINE-FORMAT.md](https://github.com/ConardLi/garden-skills/blob/main/skills/web-video-presentation/references/OUTLINE-FORMAT.md).

How do I switch from Minimax to OpenAI TTS for audio synthesis?

Set the PRESENTATION_TTS environment variable to openai before running the synthesis command:

export PRESENTATION_TTS=openai
npm run synthesize-audio

Ensure your OPENAI_API_KEY is available in the environment. The system will route requests to OpenAI's TTS endpoint instead of the default minimax provider.

What is the purpose of Checkpoint Plan?

Checkpoint Plan is a mandatory pause after Phase 1 that forces alignment on five critical items: the finalized script, the chapter outline, the selected visual theme, the asset checklist, and the development mode (A, B, or C). This prevents costly revisions by locking creative decisions before any code is written.

Can I build multiple chapters simultaneously?

Yes. Select Mode C (Parallel/Sub-agent) at Checkpoint Plan to dispatch sub-agents that work on distinct chapters concurrently. This mode is ideal for long presentations with loosely coupled sections, though the anchor chapter (Chapter 1) must still be completed and approved first to establish the visual and narrative baseline.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →