How the Web-Video-Presentation Skill Converts Scripts into Web Presentations

The web-video-presentation skill converts narration scripts into interactive, full-screen web presentations by parsing beat-delimited content into React chapter arrays, scaffolding a themed Vite+TypeScript project, and orchestrating click-advanced steps with optional AI-synthesized audio.

The ConardLi/garden-skills repository provides an automation skill that bridges scriptwriting and video production. This web-video-presentation skill transforms static text into a standalone browser application where each script beat renders as a distinct, recordable slide designed for cinematic screen recording.

Pipeline Architecture from Script to Stage

The conversion process follows an eight-stage pipeline defined in the skill's README.md. The pipeline accepts either a raw article or a pre-written script, ultimately outputting a runnable React application.

Beat Detection and Outline Generation

The pipeline begins with content formatted according to references/SCRIPT-STYLE.md. When processing an article, the agent first converts it into a spoken-style script, then splits the script into discrete beats using triple-dash delimiters (---). Each block represents one full-screen step in the final presentation.

The agent then generates an outline.md file following the specification in references/OUTLINE-FORMAT.md. This document records the step count, estimated timing, and an "information pool" extracted from the source article, serving as the blueprint for the React application.

Project Scaffolding and Theme Injection

The scripts/scaffold.sh script initiates the technical build phase. It executes npm create vite@latest with the React+TypeScript template to generate a fresh project structure, then copies boilerplate files from the templates/ directory.

To scaffold a new presentation with the "paper-press" theme:

bash skills/web-video-presentation/scripts/scaffold.sh ./my-talk --theme=paper-press
cd my-talk
npm install
npm run dev               # starts Vite dev server at http://localhost:5174

Key template files include:

The scaffold automatically injects the selected visual theme (such as "paper-press" or "warm-keynote") into the new project, establishing the color, typography, and layout tokens that drive the presentation's visual language.

Structuring Content with Chapter Registries

Once scaffolded, the project organizes content through a registry pattern that connects React components to their corresponding narration arrays.

The Narration Array Pattern

Each chapter stores its script beats in a dedicated narrations.ts file within src/chapters/<chapter-id>/. This module exports a simple string array where each index corresponds to a presentation step:

// src/chapters/02-demo/narrations.ts
export const narrations = [
  "今天我们来聊聊如何把脚本变成视频。",
  "第一步,拆分脚本成每个段落的节拍。",
  "接着,把每段写进这个数组。",
];

Silent steps—array entries without text—are automatically omitted during the audio extraction phase, allowing presenters to create pause beats in the visual flow.

Registry Configuration

The src/registry/chapters.ts file (copied from templates/src/registry/chapters.ts during scaffolding) imports each chapter's component and narration array, then exports a CHAPTERS constant:

// src/registry/chapters.ts
import DemoChapter from "../chapters/02-demo/Demo";
import { narrations as demoNarrations } from "../chapters/02-demo/narrations";

export const CHAPTERS = [
  {
    id: "demo",
    title: "Demo 章节",
    narrations: demoNarrations,
    Component: DemoChapter,
  },
];

This registry drives both the runtime rendering and the static audio manifest generation.

Audio Extraction and Synthesis Pipeline

The skill provides tooling to convert text narrations into synchronized audio files, creating a cinematic viewing experience.

Generating the Audio Manifest

The scripts/extract-narrations.ts utility reads the chapter registry and dynamically imports every narrations.ts file (lines 9-23). It flattens these arrays into a machine-readable audio-segments.json manifest describing each step's metadata:


# Extract narration segments (produces audio-segments.json)

npm run extract-narrations -- --print

The resulting JSON maps text content to filesystem paths:

{
  "chapter": "demo",
  "step": 0,
  "text": "今天我们来聊聊如何把脚本变成视频。",
  "audio": "public/audio/demo/0.mp3"
}

TTS Integration and File Generation

If audio synthesis is required, the npm run synthesize-audio command executes scripts/synthesize-audio.sh. This shell script iterates through audio-segments.json and calls pluggable providers located in scripts/tts-providers/, supporting engines like MiniMax and OpenAI.

Generated files are written to public/audio/<chapter>/<step>.mp3 (lines 15-20 of extract-narrations.ts), matching the paths referenced in the JSON manifest. The runtime's useAudioPlayer.ts hook consumes these files during playback.

Runtime Presentation Engine

The generated application functions as a fixed 16:9 stage that advances through content via user interaction or automated playback.

Step Management with useStepper

The src/hooks/useStepper.ts hook manages the global step cursor, persisting the current position to localStorage so presenters can resume where they left off. It exposes nextStep() and prevStep() functions that respond to both click events and keyboard shortcuts (arrow keys).

This hook ensures that navigation state remains consistent across page reloads while allowing seamless transitions between chapters.

Stage Rendering and Interaction

The src/components/Stage.tsx component renders the active chapter's React component for the current step index. It implements chromeless UI behavior—hiding controls until the user hovers—ensuring clean screen recordings without distracting interface elements.

The stage maintains a rigid 1920×1080 resolution scaling via useStageScale, guaranteeing consistent output dimensions regardless of browser window size. When audio assets exist, useAudioPlayer.ts synchronizes playback with step transitions, automatically advancing to the next slide when narration completes if auto-mode is enabled.

Recording and Playback Modes

The presentation supports multiple runtime modes controlled via URL query parameters:

  • ?audio=1 – Enables audio playback if MP3 files exist
  • ?auto=1 – Activates automatic advancement, ideal for hands-free screen recording

To capture the final video, developers run the Vite dev server (npm run dev), open the browser at the specified port (typically http://localhost:5174), and record the 1920×1080 window using any screen recording software.

Summary

Frequently Asked Questions

How does the skill split a script into presentation steps?

According to the references/SCRIPT-STYLE.md specification in the ConardLi/garden-skills repository, the skill uses triple-dash delimiters (---) to partition scripts into discrete beats. Each delimiter-separated block becomes one step in the presentation outline and one entry in the chapter's narrations.ts array.

Can the presentations include synchronized audio narration?

Yes. The scripts/extract-narrations.ts tool generates an audio-segments.json manifest from the registry, and scripts/synthesize-audio.sh can generate MP3 files using TTS providers like MiniMax or OpenAI. The useAudioPlayer.ts hook synchronizes playback with slide transitions during runtime.

What technology stack powers the generated presentations?

The scaffolded projects use Vite for bundling, React with TypeScript for components, and CSS theme tokens for styling. The runtime includes custom hooks like useStepper.ts and useStageScale for navigation and responsive scaling, as implemented in the templates/src/ directory.

How do I record the final presentation video?

Open the development server in a browser window sized to 1920×1080 and use any screen recording software to capture the viewport. Append ?auto=1 to the URL for hands-free automatic advancement, or ?audio=1 to include synchronized narration in the recording.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →