How to Use video-downloader Skill for Video Content Preparation: A Complete Guide

Use the video-downloader skill to download videos, extract subtitles or generate transcripts via ASR, then feed the resulting text file into cangjie-skill for methodology extraction and skill generation.

The video-downloader skill in the kangarooking/cangjie-skill repository solves a critical bottleneck in video-based knowledge work: converting audiovisual content into machine-readable text. Without this preprocessing step, the cangjie-skill pipeline cannot process videos because its entire RIA-TV++ methodology operates on plain text, not media files.

Why Video Transcription Is Required

The cangjie-skill pipeline was architected for books and long-form text, not raw video files. According to the repository's SKILL.md definition, video or podcast sessions must first be processed by a video-downloader tool to obtain transcription before any extraction occurs【https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md#L44-L45】.

This design decision reflects a separation of concerns:

  • video-downloader handles the messy domain of video codecs, streaming protocols, subtitle formats, and speech-to-text engines
  • cangjie-skill focuses exclusively on semantic analysis, methodology extraction, and skill synthesis

The README.md explicitly recommends this pairing:

"如果要蒸馏视频内容,建议搭配 video-downloader skill 一起使用:先用它下载视频、提取字幕/音频转写和关键素材,再把得到的文本内容交给 cangjie-skill 做方法论抽取、skill 化和压力测试。"【https://github.com/kangarooking/cangjie-skill/blob/main/README.md#L32-L33】

What video-downloader Actually Does

When invoked with a video URL (B站, YouTube, or other platforms), the skill performs three operations:

  1. Downloads the video file to local storage
  2. Extracts embedded subtitles if available, or runs automatic speech recognition (ASR) if not
  3. Outputs a UTF-8 encoded .srt or .txt file and returns its filesystem path

This output path becomes the "内容文本来源" (content text source) for the downstream cangjie-skill pipeline.

The Complete Workflow: video-downloader → Transcript → cangjie-skill

Step 1: Generate the Transcript

Use the video-downloader skill via Claude Code CLI:

claude run video-downloader \
  --url "https://www.bilibili.com/video/BV1x5411Y7gD" \
  --output-dir "./tmp/video1"

The skill prints the transcript path upon completion:

TRANSCRIPT_PATH=./tmp/video1/video1.srt

For YouTube content, the invocation pattern remains identical:

claude run video-downloader \
  --url "https://youtu.be/7T0jLFXE5Xc" \
  --output-dir "./tmp/yt_video"

Step 2: Feed Transcript to cangjie-skill

Pass the transcript path as the --text-path argument along with metadata:

claude run cangjie-skill \
  --text-path "./tmp/video1/video1.srt" \
  --metadata '{"title":"AI for Everyone","author":"Andrew Ng","date":"2023-04-15"}'

The skill automatically creates a folder structure under books/ai-for-everyone/ containing:

Step 3: Automated Integration (Python)

For batch processing or CI/CD pipelines, wrap both skills in Python:

import json
import subprocess

# Phase 1: Download and transcribe

result = subprocess.check_output([
    "claude", "run", "video-downloader",
    "--url", "https://youtu.be/7T0jLFXE5Xc",
    "--output-dir", "./tmp/yt_video"
])
transcript_path = json.loads(result)["TRANSCRIPT_PATH"]

# Phase 2: Extract skills from transcript

subprocess.run([
    "claude", "run", "cangjie-skill",
    "--text-path", transcript_path,
    "--metadata", json.dumps({
        "title": "AI for Everyone",
        "author": "Andrew Ng",
        "date": "2023-04-15"
    })
])

Key Source Files and Their Roles

File Purpose Location
README.md Documents the video-downloader → cangjie-skill integration pattern README.md
SKILL.md Formal specification requiring transcript preprocessing for video sources SKILL.md
methodology/00-overview.md Describes the 7-stage RIA-TV++ pipeline that consumes transcripts methodology/00-overview.md
extractors/ Prompt templates for parallel extraction of frameworks, principles, cases, counter-examples, and glossary terms extractors/
templates/ Markdown templates for generated outputs including SKILL.md.template templates/
video-downloader (external) Implements the download, subtitle extraction, and ASR transcription logic video-downloader

Handling Edge Cases

No embedded subtitles: The video-downloader skill automatically falls back to ASR transcription. Expect longer processing times and verify technical terminology in the output.

Multi-part videos: Process each URL sequentially, then concatenate transcripts before cangjie-skill ingestion, or run separate skill extraction passes and merge outputs manually.

Non-ASCII content: Both skills enforce UTF-8 encoding throughout the pipeline. Ensure your terminal and filesystem support this.

Summary

  • video-downloader is a mandatory preprocessing step for any video source in the cangjie-skill ecosystem
  • The skill returns a filesystem path to a .srt or .txt transcript, never raw video data
  • cangjie-skill's SKILL.md formalizes this dependency; the README.md provides operational guidance
  • The complete chain—video-downloader → transcript → cangjie-skill—enables the full RIA-TV++ methodology on video content
  • Resulting artifacts include structured SKILL.md files and testable prompts ready for the darwin-skill evolution stage

Frequently Asked Questions

What video platforms does video-downloader support?

The skill supports B站 (bilibili.com) and YouTube URLs explicitly, with architecture designed for extensibility to other platforms. The underlying download logic handles streaming protocols and rate limiting automatically.

Can I use cangjie-skill directly on a video file without video-downloader?

No. According to the SKILL.md specification【https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md#L44-L45】, video sources must be preprocessed to obtain a plain-text transcript. The cangjie-skill pipeline has no video parsing capabilities.

What format should the transcript be in for cangjie-skill to accept it?

The video-downloader skill outputs either SRT subtitle files (with timing codes) or plain TXT files. Both formats are valid --text-path inputs. The cangjie-skill extractors parse content regardless of timestamp presence.

How long does the video-downloader transcription process take?

Embedded subtitle extraction completes in seconds. ASR transcription scales linearly with video duration—typically 0.5-2x real-time depending on hardware and whether GPU acceleration is available for the speech recognition model.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →