How to Convert Articles to Narration Scripts with Garden Skills: A Complete Guide

Garden Skills provides a method-driven agent framework that transforms raw articles into production-ready narration scripts through the web-video-presentation skill, enforcing style guidelines and validation checkpoints at each phase.

The web-video-presentation skill in the ConardLi/garden-skills repository implements a structured pipeline to convert articles into cinematic video presentations. This skill automates the transformation of text into script.md files that serve as the foundation for visual outlines and optional audio synthesis.

Understanding the web-video-presentation Skill Architecture

The skill operates as a phase-driven agent defined in /skills/web-video-presentation/SKILL.md. It treats article conversion as a formal production pipeline rather than simple text transformation. The repository organizes functionality into distinct components:

  • SCRIPT-STYLE.md enforces sentence-level density rules and language preservation standards
  • OUTLINE-FORMAT.md governs how narration beats map to visual scenes
  • AUDIO.md provides TTS provider contracts for Minimax, OpenAI, ElevenLabs, and Azure
  • manifest.json drives the skill's execution parameters and stage definitions

The Article-to-Script Conversion Pipeline

Phase 1.2: Generating the Narration Script

The conversion process begins when the skill ingests source material—whether plain text, Markdown, or PDF-derived content—and produces a platform-agnostic script.md. According to /skills/web-video-presentation/references/SCRIPT-STYLE.md, the generated script must adhere to strict formatting rules regarding sentence density and structural consistency.

Key generation parameters include:

  • Sentence-level density controls preventing overly complex narration blocks
  • Language preservation maintaining the source article's linguistic characteristics
  • Beat-by-beat structure defining precise narration timing for downstream visual mapping

Validation and Checkpoint Requirements

Before the pipeline advances, the skill executes three layers of self-validation as documented in /skills/web-video-presentation/SKILL.md lines 87-94:

  1. Format validation ensuring script.md conforms to structural templates
  2. Tone verification checking consistency with the target presentation style
  3. Read-aloud testing verifying natural speech patterns and timing

The system then enforces a hard checkpoint requiring explicit user confirmation of the script before proceeding to outline generation. This prevents downstream errors from propagating into visual production.

From Script to Outline and Audio

Phase 2: Outline Generation

Once validated, script.md feeds into the outline generation phase. The skill parses each narration beat to create outline.md, which maps spoken content to full-screen visual steps. This process respects the step size, timing, and scene boundaries defined in /skills/web-video-presentation/references/OUTLINE-FORMAT.md.

The outline serves as the bridge between audio narration and visual presentation, ensuring that every spoken beat corresponds to a specific visual element or transition.

Optional Audio Synthesis

The skill supports automated narration generation through the tts-providers implementation documented in /skills/web-video-presentation/references/AUDIO.md lines 13-22. Available providers include:

  • Minimax (built-in support)
  • OpenAI (built-in support)
  • ElevenLabs (via ready-made integration snippets)
  • Azure Speech Services (via configuration templates)

Audio synthesis occurs after outline approval, generating synchronized audio files for each narration beat defined in the script.

Implementing the Conversion Workflow

Project Scaffolding

Initialize a new presentation workspace using the scaffold script:

bash skills/web-video-presentation/scripts/scaffold.sh ./my-presentation --theme=paper-press

This command creates a Vite + React + TypeScript project structure including script.md, outline.md, and theme assets as referenced in /skills/web-video-presentation/scripts/scaffold.sh lines 10-12.

Executing the Conversion

Convert an existing article into a narration script:

garden run web-video-presentation --input article.md --stage script

Inspect the generated output:

cat my-presentation/script.md

Generating Downstream Assets

Create the visual outline from the approved script:

garden run web-video-presentation --input script.md --stage outline

Optionally synthesize narration audio:

bash my-presentation/scripts/synthesize-audio.sh

Summary

  • Garden Skills automates article-to-script conversion through the web-video-presentation skill, enforcing strict formatting via SCRIPT-STYLE.md.
  • The pipeline requires three-layer validation (format, tone, read-aloud) and a hard user checkpoint before accepting the generated script.md.
  • Validated scripts automatically generate visual outlines mapped to full-screen presentation steps according to OUTLINE-FORMAT.md.
  • Optional TTS synthesis supports Minimax, OpenAI, ElevenLabs, and Azure providers through standardized audio contracts.
  • The scaffold script bootstraps complete Vite-based projects ready for screen recording and cinematic presentation.

Frequently Asked Questions

What input formats does the web-video-presentation skill support?

The skill accepts multiple input formats including plain text, Markdown, and PDF-derived content. According to /skills/web-video-presentation/SKILL.md, the parser normalizes these inputs during Phase 1.2 to produce a consistent script.md output regardless of the source format.

How does the skill ensure narration quality before proceeding?

The system implements three validation layers documented in /skills/web-video-presentation/SKILL.md lines 87-94: format verification against SCRIPT-STYLE.md, tone consistency checks, and read-aloud testing for natural speech flow. After these automated checks, a hard checkpoint requires explicit user confirmation, preventing low-quality scripts from advancing to outline generation.

Can I customize the TTS provider for audio generation?

Yes. While the skill includes built-in support for Minimax and OpenAI, /skills/web-video-presentation/references/AUDIO.md provides integration snippets for ElevenLabs and Azure Speech Services. You can configure the desired provider through the skill's manifest or environment variables before running the synthesis scripts.

What is the relationship between script.md and outline.md?

script.md contains the narration beats—the spoken content and timing—while outline.md maps those beats to visual scenes. As specified in /skills/web-video-presentation/references/OUTLINE-FORMAT.md, the outline translates each script segment into full-screen visual steps with precise timing boundaries, creating the storyboard for the final presentation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →