Architecture Principles Behind the Web-Video-Presentation Skill: Cinematic Design for AI Decks
The web-video-presentation skill implements a method-driven architecture that treats web presentations as cinematic video production surfaces, combining a fixed 16:9 coordinate system, global step cursor, script-driven content beats, and hard collaboration checkpoints to ensure deterministic, high-fidelity output.
The architecture principles behind the web-video-presentation skill in the ConardLi/garden-skills repository redefine automated presentation generation by embedding cinematic production constraints directly into the codebase. Rather than assembling traditional slide decks, this system orchestrates a deterministic workflow governed by markdown contracts, semantic design tokens, and provider-agnostic audio synthesis that guarantees broadcast-quality visual consistency.
Fixed-Stage Cinematic Canvas
At the foundation lies a fixed 16:9 stage that establishes a stable 1920 × 1080 coordinate system for all content authoring.
This coordinate system guarantees predictable layout behavior specifically optimized for screen-recording workflows. Content scales responsively to the viewport while maintaining the exact aspect ratio, ensuring that recordings captured at any resolution remain compositionally consistent. As documented in skills/web-video-presentation/README.md, this fixed stage eliminates the variable positioning issues common in responsive slide decks, treating the browser window as a cinematic video production surface.
Global Step Cursor and Navigation Model
The skill implements a one global step cursor that tracks position through a (chapter, step) tuple, persisting this state locally so the presentation always knows the current beat.
Navigation advances via click or keyboard input, but the architecture strictly enforces the one step = one idea principle. Each beat maps to a full-screen scene without bullet-point accumulations, keeping every visual slice focused and cinematic. This constraint, enforced through the outline format specified in references/OUTLINE-FORMAT.md, prevents the cognitive overload typical of traditional presentation software.
Script-Driven Content Architecture
Presentation structure originates from narration scripts rather than slide layouts, embodying a script-driven structure that ensures tight audio-visual rhythm.
The source code in templates/scripts/extract-narrations.ts demonstrates how the system parses script beats separated by --- delimiters:
// file: templates/scripts/extract-narrations.ts
import { readFileSync } from 'fs';
const script = readFileSync('script.md', 'utf‑8');
const beats = script.split('\n---\n'); // each beat separated by "---"
console.log(JSON.stringify(beats, null, 2));
These extracted beats directly drive the chapter and step outline, ensuring that visual transitions align precisely with narration timing.
Visual Design Philosophy
Motion-First Mindset
Every scene must contain a moving visual anchor; static paragraphs are treated as an architectural "smell" that should be replaced with motion. This motion-first mindset forces the agent to generate dynamic visual content rather than relying on text-heavy slides, as codified in references/CHAPTER-CRAFT.md.
Hidden Chrome
To maintain clean recordings, the skill implements hidden chrome—UI controls appear only on hover. This keeps the presentation surface distraction-free during video capture while preserving accessibility for live navigation.
Token-Based Theming System
Visual decisions flow through a theme-token architecture defined in semantic design tokens rather than hardcoded values.
The contract lives in references/THEMES.md, with concrete implementations residing in themes/*/tokens.css and themes/*/theme.json. Changing a theme swaps the entire visual language without requiring content modifications. The system ships with 23 built-in themes, accessible via the scaffold command:
# List all available themes
bash templates/scripts/scaffold.sh --list-themes
# Scaffold with specific theme
bash templates/scripts/scaffold.sh ./my-presentation --theme=paper-press
Pluggable TTS Architecture
Audio synthesis operates through a provider-agnostic system where additional backends integrate by dropping shell scripts into templates/scripts/tts-providers/.
The skill includes two built-in providers: MiniMax (mmx-cli) and OpenAI TTS. The generic runner templates/scripts/synthesize-audio.sh dispatches to any executable in the providers directory:
# Execute provider-agnostic synthesis
bash templates/scripts/synthesize-audio.sh ./my-presentation/outlines/outline.md
Extending support requires only a new shell script following the interface convention. For example, adding ElevenLabs support involves creating templates/scripts/tts-providers/elevenlabs.sh:
#!/usr/bin/env bash
# Usage: ./elevenlabs.sh "Hello world" output.wav
curl -X POST "https://api.elevenlabs.io/v1/text-to-speech/..." \
-H "xi‑api‑key: $ELEVENLABS_API_KEY" \
-d "{\"text\":\"$1\"}" --output "$2"
Hard Collaboration Checkpoints
The architecture enforces hard collaboration checkpoints through the skill contract defined in SKILL.md, forcing the agent to pause at three explicit stages: (A1) script and theme alignment, (A2) outline approval, and (B) optional audio synthesis.
These checkpoints prevent silent shortcuts and maintain human-in-the-loop verification before computationally expensive operations or irreversible structural commitments are made.
Contract-Driven Development
The workflow relies on reference-driven contracts—a set of markdown specifications in the references/ directory that codify agent behavior:
references/PRINCIPLES.md– Core architectural constraintsreferences/CHAPTER-CRAFT.md– Rules for chapter implementation and motion requirementsreferences/OUTLINE-FORMAT.md– Required outline structure driving step-by-step presentationreferences/AUDIO.md– Narration synthesis workflow specifications
These declarative contracts make the workflow deterministic and reproducible across different content inputs.
Implementation Scaffold
The skill generates a Vite + React + TypeScript project through templates/scripts/scaffold.sh, exposing reusable stage primitives and recording guidance. This scaffold implements the fixed-stage coordinate system and global cursor navigation, providing the runtime environment for the architectural principles to execute.
Summary
- The web-video-presentation skill uses a fixed 1920 × 1080 coordinate system to ensure predictable screen-recording layouts.
- A global (chapter, step) cursor enforces the "one step, one idea" principle across full-screen cinematic scenes.
- Script-driven extraction parses narration beats from markdown files to drive visual structure.
- Theme-token architecture enables complete visual language swaps via semantic tokens without content changes.
- Pluggable TTS supports MiniMax and OpenAI out of the box, with extensibility through shell scripts in
tts-providers/. - Hard checkpoints (A1, A2, B) enforce human approval before major workflow transitions.
- Reference contracts in the
references/directory codify deterministic agent behavior.
Frequently Asked Questions
How does the fixed 16:9 stage handle different screen resolutions?
The stage maintains a 1920 × 1080 authoring coordinate system that scales responsively to the viewport while preserving aspect ratio. According to skills/web-video-presentation/README.md, this guarantees that content positioning remains predictable during screen recording regardless of the capture resolution, treating the browser as a cinematic production surface.
What is the purpose of the hard collaboration checkpoints?
The checkpoints force the AI agent to pause for human verification at three critical stages: A1 (script and theme alignment), A2 (outline approval), and B (audio synthesis). As specified in SKILL.md, these hard stops prevent silent errors and ensure the human remains in the loop before expensive computational steps or structural commitments.
How do I add a new text-to-speech provider to the skill?
Create a shell script in templates/scripts/tts-providers/ that accepts text and output file arguments. The generic runner synthesize-audio.sh automatically detects and dispatches to any executable in that directory, making the system provider-agnostic without modifying core logic.
Where are the design tokens and theming rules defined?
Semantic design tokens are specified in references/THEMES.md, with concrete implementations residing in themes/*/tokens.css and themes/*/theme.json. Changing themes swaps the entire visual language without modifying presentation content, following the theme-token architecture principle implemented in the scaffold templates.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →