How to Handle Non-Book Content Like Video Timestamps in cangjie-skill
Strip timestamp lines from subtitle files before feeding them to the cangjie pipeline to prevent the RIA-TV++ extractors from treating temporal markers as semantic content.
The kangarooking/cangjie-skill repository transforms high-value textual material—from books to video transcripts—into structured AI Skills. When processing video content, timestamps in subtitle files (e.g., 00:01:23 --> 00:01:26) must be removed before ingestion, as the pipeline's semantic extractors interpret these markers as data noise rather than content.
Why Timestamps Break the cangjie Pipeline
The cangjie-skill pipeline relies on RIA-TV++ extractors located in the extractors/ directory to identify semantic patterns like frameworks, principles, and case studies. These extractors expect clean prose, not temporal metadata.
Timestamps cause three specific failures:
- Pattern Interruption: The extractors scan for conceptual tokens. A timestamp line such as
00:01:23 --> 00:01:26contains no semantic value and breaks the context window for downstream analysis. - Triple Verification Failure: The verification stage requires at least two independent citations from meaningful text. Timestamps cannot serve as valid citations, causing extractor candidates to be rejected.
- Zettelkasten Linking Degradation: The linking heuristics in the pipeline compare pure textual tokens to build knowledge graphs. Punctuation-heavy timestamps reduce match quality between related concepts.
Preprocessing Video Subtitles for cangjie-skill
Processing non-book content requires a three-phase workflow before invoking the main cangjie.py script.
1. Extract Transcripts with video-downloader
First, convert video content to text using the companion video-downloader skill referenced in the repository's README.md. This utility downloads the source video and either extracts embedded subtitle tracks (.srt, .vtt) or runs speech-to-text to generate a raw transcript file.
2. Clean and Strip Timestamps
Run a preprocessing script to remove timestamp lines and formatting artifacts. The pipeline requires plain text input to generate valid skill artifacts like BOOK_OVERVIEW.md and INDEX.md.
3. Execute the cangjie Pipeline
Feed the cleaned transcript to the main entry point:
python cangjie.py --input ./clean/video.txt --output ./skills
This command runs the Adler overview, parallel extractors, triple verification, RIA++ construction, Zettelkasten linking, and pressure-testing to emit final artifacts including DIGEST.md and individual */SKILL.md files.
Python Utility for Timestamp Removal
Create a local helper script to sanitize subtitle files. This script removes lines matching the SRT/VTT timestamp pattern (HH:MM:SS --> HH:MM:SS) and strips index numbers and empty lines.
# clean_subtitles.py – removes SRT/VTT timestamps and writes plain text
import re, sys, pathlib
TIMESTAMP_RE = re.compile(r'^\s*\d{2}:\d{2}:\d{2}(?:[.,]\d+)?\s*-->?\s*\d{2}:\d{2}:\d{2}(?:[.,]\d+)?\s*$')
def clean_subtitle_file(src_path: pathlib.Path, dst_path: pathlib.Path) -> None:
"""Read a subtitle file (SRT/VTT) and write a timestamp‑free plain‑text version."""
with src_path.open(encoding="utf-8") as src, dst_path.open("w", encoding="utf-8") as dst:
for line in src:
if TIMESTAMP_RE.match(line):
continue # skip timestamps
if line.strip().isdigit():
continue # skip index numbers (optional)
if line.strip() == "":
continue # skip empty lines
dst.write(line.strip() + " ")
if __name__ == "__main__":
if len(sys.argv) != 3:
sys.exit("Usage: python clean_subtitles.py <input.srt> <output.txt>")
clean_subtitle_file(pathlib.Path(sys.argv[1]), pathlib.Path(sys.argv[2]))
Typical usage:
# Download video & get subtitles
python video_downloader.py --url "https://www.youtube.com/watch?v=XYZ" --out-dir ./raw
# Clean the subtitle file
python clean_subtitles.py ./raw/video.srt ./clean/video.txt
# Run cangjie-skill on the cleaned transcript
python cangjie.py --input ./clean/video.txt --output ./skills
Optional: Preserving Timestamp References
If traceability is required, store the original subtitle file alongside the cleaned version. You can map generated skill entries back to specific video timestamps by storing line range metadata in a JSON sidecar:
{
"skill_id": "example-skill",
"source": {
"subtitle_file": "video.srt",
"line_range": "45-47"
}
}
The generated SKILL.md files can reference this mapping in their source fields, allowing users to locate the exact moment in the original video while keeping the pipeline's input free of temporal noise.
Summary
- Timestamps must be stripped from video subtitles before processing with
cangjie.pyto avoid RIA-TV++ extractor failures. - Use the companion
video-downloaderskill to extract.srtor.vttfiles from video sources. - Run a Python preprocessing script (like
clean_subtitles.py) to remove timestamp patterns, index numbers, and empty lines. - Feed only the cleaned plain text to the cangjie pipeline to ensure valid
BOOK_OVERVIEW.md,INDEX.md, andSKILL.mdgeneration. - Retain original subtitle files for reference mapping if source traceability is required.
Frequently Asked Questions
What happens if I don't remove timestamps from video transcripts?
The RIA-TV++ extractors will treat timestamp strings as content tokens, causing the triple verification step to fail when it cannot find meaningful semantic citations. This results in empty or incomplete skill artifacts because the extractors cannot validate concepts against temporal metadata.
Can cangjie-skill process audio podcasts the same way as videos?
Yes. Audio content requires the same preprocessing: transcribe the audio to text using a speech-to-text engine, strip any generated timestamps or speaker labels, and process the resulting plain text through cangjie.py. The pipeline handles any long-form text regardless of original media type.
Where does the cleaned text get stored in the cangjie output structure?
The cleaned input text is not stored in the output; rather, the pipeline generates structured artifacts in the --output directory, including BOOK_OVERVIEW.md (high-level summary), INDEX.md (searchable master list), DIGEST.md (condensed takeaways), and individual SKILL.md files in subdirectories as defined in the repository's SKILL.md schema specification.
Is there an automated way to handle timestamp removal in CI/CD?
Yes. Integrate the clean_subtitles.py script into your CI workflow before the cangjie execution step. The repository's .github/workflows/ directory (containing examples like update-star-history.yml) demonstrates the project's automation patterns, which you can extend to include subtitle preprocessing as a prerequisite job before running the main skill generation pipeline.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →