Web Presentation Workflow Phases: A 4-Step Pipeline for Creating Click-Driven 16:9 Demos

The web-video-presentation skill implements a four-phase pipeline with three mandatory checkpoints that transforms raw articles into screen-ready web demos using a fixed 1920×1080 stage and theme-driven architecture.

This workflow, defined in the ConardLi/garden-skills repository, provides a reproducible method for converting written content into interactive, click-driven presentations. By enforcing hard checkpoints between phases, the workflow ensures script-theme alignment and visual consistency before generating the final screen-recording artifact.

The 4-Phase Web Presentation Workflow

According to the README.md in skills/web-video-presentation/, the pipeline is divided into four distinct phases with three hard checkpoints that force user confirmation.

Phase 1: Content Preparation and Scripting

The initial phase breaks down into three sub-steps that handle content transformation:

Phase 1.1 – Identify Input: Determine the source material, whether it is a raw article or an existing script.

Phase 1.2 – Article → Narration Script: Convert the article into a narration-friendly script. This step triggers Checkpoint A1, where you must review the generated script, agree on a theme, and draft an asset plan before proceeding.

Phase 1.3 – Script + Article → Outline: Combine the script and source article to produce an outline.md file that drives the visual flow. This requires Checkpoint A2, where you approve the outline and decide on the development mode (prototype or full build).

Phase 2: Build the Presentation

In this phase, the system scaffolds a Vite + React + TypeScript project using skills/web-video-presentation/scripts/scaffold.sh. The selected theme is applied immediately, and each step of the outline is implemented as a full-screen scene within the fixed 1920×1080 coordinate system.

The phase concludes with Checkpoint B, which asks whether to synthesize audio for the presentation. This optional step allows you to proceed with or without generated narration.

Phase 3: Optional Audio Synthesis

If enabled at Checkpoint B, this phase runs the pluggable TTS (Text-to-Speech) pipeline via templates/scripts/synthesize-audio.sh. The script provider architecture in templates/scripts/tts-providers/ allows swapping between different TTS backends (such as minimax.sh or openai.sh) without modifying the core workflow.

Phase 4: Recording and Post-Production

The final phase involves recording the 1920×1080 stage using a screen-recorder, then applying post-production tweaks such as trimming and encoding. The fixed aspect ratio ensures consistent output dimensions suitable for video platforms.

Key Technical Implementation Details

The workflow enforces specific constraints that distinguish it from traditional slide-based presentations.

Fixed 16:9 Stage: All scenes are authored in a stable 1920×1080 coordinate system defined in the skill contract (SKILL.md). This guarantees a clean recording canvas regardless of the display device.

Theme-Token Architecture: Selecting a theme early (during Checkpoint A1) propagates token values for colors, typography, and motion throughout the build. The themes/ directory contains 23 pre-built theme folders, each defining these token contracts, making visual redesign inexpensive.

One-Step, One-Idea Navigation: Each click advances a single narration beat, preventing "slide-bullet" overload and maintaining tight visual narrative control as outlined in references/PRINCIPLES.md.

Getting Started: Scaffold Your First Presentation

To execute the workflow phases locally, use the CLI helpers provided in the repository:


# Scaffold a new presentation with a specific theme (e.g., "paper-press")

bash skills/web-video-presentation/scripts/scaffold.sh ./my-presentation --theme=paper-press

# List all 23 available themes

bash skills/web-video-presentation/scripts/scaffold.sh --list-themes

# Start the development server

cd my-presentation && npm install && npm run dev

After completing Phase 2 (Build), optionally run Phase 3 audio synthesis:


# Generate narration audio using the default Minimax provider

bash skills/web-video-presentation/templates/scripts/synthesize-audio.sh ./my-presentation

Core Files and Architecture

The following files implement the phased workflow in the ConardLi/garden-skills repository:

Summary

  • The web-video-presentation skill uses a four-phase pipeline (Content Preparation, Build, Audio Synthesis, Recording) with three mandatory checkpoints (A1, A2, and B) to ensure quality control.
  • Each phase produces specific artifacts: narration scripts, outline.md, Vite/React code, optional audio files, and final screen recordings.
  • The fixed 1920×1080 stage and theme-token architecture ensure consistent visual output across all presentations.
  • Audio synthesis is pluggable and optional, supporting multiple TTS providers through the tts-providers/ directory structure.
  • The workflow is implemented in the ConardLi/garden-skills repository using shell scripts, React components, and markdown-based outline files.

Frequently Asked Questions

What are the hard checkpoints in the web presentation workflow?

The workflow enforces three hard checkpoints that pause execution for user confirmation. Checkpoint A1 occurs after converting an article to a narration script, requiring theme selection and asset planning. Checkpoint A2 follows outline generation, requiring approval of the visual flow and development mode decision. Checkpoint B appears after building the presentation, asking whether to proceed with optional audio synthesis.

How does the theme system work in the web-video-presentation skill?

The theme system uses a token-based architecture where selecting a theme during Checkpoint A1 propagates values for colors, typography, and motion throughout the entire build. The themes/ directory contains 23 pre-built themes, and the scaffold.sh script applies these tokens when creating the Vite + React + TypeScript project, making global style changes inexpensive.

Can I use different text-to-speech providers for audio generation?

Yes, the audio synthesis phase supports pluggable TTS backends. The templates/scripts/synthesize-audio.sh script acts as a provider-agnostic runner, while individual providers (such as minimax.sh or openai.sh) reside in the templates/scripts/tts-providers/ directory. You can swap providers by selecting different scripts or adding new ones without modifying the core workflow logic.

What resolution and aspect ratio does the workflow target?

All scenes are authored in a fixed 16:9 stage with a 1920×1080 coordinate system. This fixed canvas, defined in the skill contract, guarantees consistent recording dimensions suitable for video platforms and prevents responsive layout issues during screen recording.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →