How to Use video-downloader Skill for Video Content Preparation: A Complete Guide
Use the video-downloader skill to download videos, extract subtitles or generate transcripts via ASR, then feed the resulting text file into cangjie-skill for methodology extraction and skill generation.
The video-downloader skill in the kangarooking/cangjie-skill repository solves a critical bottleneck in video-based knowledge work: converting audiovisual content into machine-readable text. Without this preprocessing step, the cangjie-skill pipeline cannot process videos because its entire RIA-TV++ methodology operates on plain text, not media files.
Why Video Transcription Is Required
The cangjie-skill pipeline was architected for books and long-form text, not raw video files. According to the repository's SKILL.md definition, video or podcast sessions must first be processed by a video-downloader tool to obtain transcription before any extraction occurs【https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md#L44-L45】.
This design decision reflects a separation of concerns:
- video-downloader handles the messy domain of video codecs, streaming protocols, subtitle formats, and speech-to-text engines
- cangjie-skill focuses exclusively on semantic analysis, methodology extraction, and skill synthesis
The README.md explicitly recommends this pairing:
"如果要蒸馏视频内容,建议搭配 video-downloader skill 一起使用:先用它下载视频、提取字幕/音频转写和关键素材,再把得到的文本内容交给 cangjie-skill 做方法论抽取、skill 化和压力测试。"【https://github.com/kangarooking/cangjie-skill/blob/main/README.md#L32-L33】
What video-downloader Actually Does
When invoked with a video URL (B站, YouTube, or other platforms), the skill performs three operations:
- Downloads the video file to local storage
- Extracts embedded subtitles if available, or runs automatic speech recognition (ASR) if not
- Outputs a UTF-8 encoded
.srtor.txtfile and returns its filesystem path
This output path becomes the "内容文本来源" (content text source) for the downstream cangjie-skill pipeline.
The Complete Workflow: video-downloader → Transcript → cangjie-skill
Step 1: Generate the Transcript
Use the video-downloader skill via Claude Code CLI:
claude run video-downloader \
--url "https://www.bilibili.com/video/BV1x5411Y7gD" \
--output-dir "./tmp/video1"
The skill prints the transcript path upon completion:
TRANSCRIPT_PATH=./tmp/video1/video1.srt
For YouTube content, the invocation pattern remains identical:
claude run video-downloader \
--url "https://youtu.be/7T0jLFXE5Xc" \
--output-dir "./tmp/yt_video"
Step 2: Feed Transcript to cangjie-skill
Pass the transcript path as the --text-path argument along with metadata:
claude run cangjie-skill \
--text-path "./tmp/video1/video1.srt" \
--metadata '{"title":"AI for Everyone","author":"Andrew Ng","date":"2023-04-15"}'
The skill automatically creates a folder structure under books/ai-for-everyone/ containing:
BOOK_OVERVIEW.md— source analysis and structural map- Multiple
SKILL.mdfiles — one per extracted methodological unit test-prompts.jsonandtest-results.md— pressure testing artifactsINDEX.mdandGLOSSARY.md— navigation and terminologyDIGEST.md— read-only executive summary
Step 3: Automated Integration (Python)
For batch processing or CI/CD pipelines, wrap both skills in Python:
import json
import subprocess
# Phase 1: Download and transcribe
result = subprocess.check_output([
"claude", "run", "video-downloader",
"--url", "https://youtu.be/7T0jLFXE5Xc",
"--output-dir", "./tmp/yt_video"
])
transcript_path = json.loads(result)["TRANSCRIPT_PATH"]
# Phase 2: Extract skills from transcript
subprocess.run([
"claude", "run", "cangjie-skill",
"--text-path", transcript_path,
"--metadata", json.dumps({
"title": "AI for Everyone",
"author": "Andrew Ng",
"date": "2023-04-15"
})
])
Key Source Files and Their Roles
| File | Purpose | Location |
|---|---|---|
README.md |
Documents the video-downloader → cangjie-skill integration pattern | README.md |
SKILL.md |
Formal specification requiring transcript preprocessing for video sources | SKILL.md |
methodology/00-overview.md |
Describes the 7-stage RIA-TV++ pipeline that consumes transcripts | methodology/00-overview.md |
extractors/ |
Prompt templates for parallel extraction of frameworks, principles, cases, counter-examples, and glossary terms | extractors/ |
templates/ |
Markdown templates for generated outputs including SKILL.md.template |
templates/ |
video-downloader (external) |
Implements the download, subtitle extraction, and ASR transcription logic | video-downloader |
Handling Edge Cases
No embedded subtitles: The video-downloader skill automatically falls back to ASR transcription. Expect longer processing times and verify technical terminology in the output.
Multi-part videos: Process each URL sequentially, then concatenate transcripts before cangjie-skill ingestion, or run separate skill extraction passes and merge outputs manually.
Non-ASCII content: Both skills enforce UTF-8 encoding throughout the pipeline. Ensure your terminal and filesystem support this.
Summary
- video-downloader is a mandatory preprocessing step for any video source in the cangjie-skill ecosystem
- The skill returns a filesystem path to a
.srtor.txttranscript, never raw video data - cangjie-skill's
SKILL.mdformalizes this dependency; theREADME.mdprovides operational guidance - The complete chain—video-downloader → transcript → cangjie-skill—enables the full RIA-TV++ methodology on video content
- Resulting artifacts include structured
SKILL.mdfiles and testable prompts ready for the darwin-skill evolution stage
Frequently Asked Questions
What video platforms does video-downloader support?
The skill supports B站 (bilibili.com) and YouTube URLs explicitly, with architecture designed for extensibility to other platforms. The underlying download logic handles streaming protocols and rate limiting automatically.
Can I use cangjie-skill directly on a video file without video-downloader?
No. According to the SKILL.md specification【https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md#L44-L45】, video sources must be preprocessed to obtain a plain-text transcript. The cangjie-skill pipeline has no video parsing capabilities.
What format should the transcript be in for cangjie-skill to accept it?
The video-downloader skill outputs either SRT subtitle files (with timing codes) or plain TXT files. Both formats are valid --text-path inputs. The cangjie-skill extractors parse content regardless of timestamp presence.
How long does the video-downloader transcription process take?
Embedded subtitle extraction completes in seconds. ASR transcription scales linearly with video duration—typically 0.5-2x real-time depending on hardware and whether GPU acceleration is available for the speech recognition model.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →