How to Convert Articles to Narration Scripts with Garden Skills: A Complete Guide
Garden Skills provides a method-driven agent framework that transforms raw articles into production-ready narration scripts through the web-video-presentation skill, enforcing style guidelines and validation checkpoints at each phase.
The web-video-presentation skill in the ConardLi/garden-skills repository implements a structured pipeline to convert articles into cinematic video presentations. This skill automates the transformation of text into script.md files that serve as the foundation for visual outlines and optional audio synthesis.
Understanding the web-video-presentation Skill Architecture
The skill operates as a phase-driven agent defined in /skills/web-video-presentation/SKILL.md. It treats article conversion as a formal production pipeline rather than simple text transformation. The repository organizes functionality into distinct components:
SCRIPT-STYLE.mdenforces sentence-level density rules and language preservation standardsOUTLINE-FORMAT.mdgoverns how narration beats map to visual scenesAUDIO.mdprovides TTS provider contracts for Minimax, OpenAI, ElevenLabs, and Azuremanifest.jsondrives the skill's execution parameters and stage definitions
The Article-to-Script Conversion Pipeline
Phase 1.2: Generating the Narration Script
The conversion process begins when the skill ingests source material—whether plain text, Markdown, or PDF-derived content—and produces a platform-agnostic script.md. According to /skills/web-video-presentation/references/SCRIPT-STYLE.md, the generated script must adhere to strict formatting rules regarding sentence density and structural consistency.
Key generation parameters include:
- Sentence-level density controls preventing overly complex narration blocks
- Language preservation maintaining the source article's linguistic characteristics
- Beat-by-beat structure defining precise narration timing for downstream visual mapping
Validation and Checkpoint Requirements
Before the pipeline advances, the skill executes three layers of self-validation as documented in /skills/web-video-presentation/SKILL.md lines 87-94:
- Format validation ensuring
script.mdconforms to structural templates - Tone verification checking consistency with the target presentation style
- Read-aloud testing verifying natural speech patterns and timing
The system then enforces a hard checkpoint requiring explicit user confirmation of the script before proceeding to outline generation. This prevents downstream errors from propagating into visual production.
From Script to Outline and Audio
Phase 2: Outline Generation
Once validated, script.md feeds into the outline generation phase. The skill parses each narration beat to create outline.md, which maps spoken content to full-screen visual steps. This process respects the step size, timing, and scene boundaries defined in /skills/web-video-presentation/references/OUTLINE-FORMAT.md.
The outline serves as the bridge between audio narration and visual presentation, ensuring that every spoken beat corresponds to a specific visual element or transition.
Optional Audio Synthesis
The skill supports automated narration generation through the tts-providers implementation documented in /skills/web-video-presentation/references/AUDIO.md lines 13-22. Available providers include:
- Minimax (built-in support)
- OpenAI (built-in support)
- ElevenLabs (via ready-made integration snippets)
- Azure Speech Services (via configuration templates)
Audio synthesis occurs after outline approval, generating synchronized audio files for each narration beat defined in the script.
Implementing the Conversion Workflow
Project Scaffolding
Initialize a new presentation workspace using the scaffold script:
bash skills/web-video-presentation/scripts/scaffold.sh ./my-presentation --theme=paper-press
This command creates a Vite + React + TypeScript project structure including script.md, outline.md, and theme assets as referenced in /skills/web-video-presentation/scripts/scaffold.sh lines 10-12.
Executing the Conversion
Convert an existing article into a narration script:
garden run web-video-presentation --input article.md --stage script
Inspect the generated output:
cat my-presentation/script.md
Generating Downstream Assets
Create the visual outline from the approved script:
garden run web-video-presentation --input script.md --stage outline
Optionally synthesize narration audio:
bash my-presentation/scripts/synthesize-audio.sh
Summary
- Garden Skills automates article-to-script conversion through the
web-video-presentationskill, enforcing strict formatting viaSCRIPT-STYLE.md. - The pipeline requires three-layer validation (format, tone, read-aloud) and a hard user checkpoint before accepting the generated
script.md. - Validated scripts automatically generate visual outlines mapped to full-screen presentation steps according to
OUTLINE-FORMAT.md. - Optional TTS synthesis supports Minimax, OpenAI, ElevenLabs, and Azure providers through standardized audio contracts.
- The scaffold script bootstraps complete Vite-based projects ready for screen recording and cinematic presentation.
Frequently Asked Questions
What input formats does the web-video-presentation skill support?
The skill accepts multiple input formats including plain text, Markdown, and PDF-derived content. According to /skills/web-video-presentation/SKILL.md, the parser normalizes these inputs during Phase 1.2 to produce a consistent script.md output regardless of the source format.
How does the skill ensure narration quality before proceeding?
The system implements three validation layers documented in /skills/web-video-presentation/SKILL.md lines 87-94: format verification against SCRIPT-STYLE.md, tone consistency checks, and read-aloud testing for natural speech flow. After these automated checks, a hard checkpoint requires explicit user confirmation, preventing low-quality scripts from advancing to outline generation.
Can I customize the TTS provider for audio generation?
Yes. While the skill includes built-in support for Minimax and OpenAI, /skills/web-video-presentation/references/AUDIO.md provides integration snippets for ElevenLabs and Azure Speech Services. You can configure the desired provider through the skill's manifest or environment variables before running the synthesis scripts.
What is the relationship between script.md and outline.md?
script.md contains the narration beats—the spoken content and timing—while outline.md maps those beats to visual scenes. As specified in /skills/web-video-presentation/references/OUTLINE-FORMAT.md, the outline translates each script segment into full-screen visual steps with precise timing boundaries, creating the storyboard for the final presentation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →