# How to Handle Non-Book Content Like Video Timestamps in cangjie-skill

> Learn to handle non-book content like video timestamps in cangjie-skill. Prevent RIA-TV++ extractors from misinterpreting temporal markers by stripping timestamp lines before processing.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-07-19

---

**Strip timestamp lines from subtitle files before feeding them to the cangjie pipeline to prevent the RIA-TV++ extractors from treating temporal markers as semantic content.**

The `kangarooking/cangjie-skill` repository transforms high-value textual material—from books to video transcripts—into structured AI Skills. When processing video content, timestamps in subtitle files (e.g., `00:01:23 --> 00:01:26`) must be removed before ingestion, as the pipeline's semantic extractors interpret these markers as data noise rather than content.

## Why Timestamps Break the cangjie Pipeline

The cangjie-skill pipeline relies on **RIA-TV++ extractors** located in the `extractors/` directory to identify semantic patterns like frameworks, principles, and case studies. These extractors expect clean prose, not temporal metadata.

Timestamps cause three specific failures:

- **Pattern Interruption**: The extractors scan for conceptual tokens. A timestamp line such as `00:01:23 --> 00:01:26` contains no semantic value and breaks the context window for downstream analysis.
- **Triple Verification Failure**: The verification stage requires at least two independent citations from meaningful text. Timestamps cannot serve as valid citations, causing extractor candidates to be rejected.
- **Zettelkasten Linking Degradation**: The linking heuristics in the pipeline compare pure textual tokens to build knowledge graphs. Punctuation-heavy timestamps reduce match quality between related concepts.

## Preprocessing Video Subtitles for cangjie-skill

Processing non-book content requires a three-phase workflow before invoking the main [`cangjie.py`](https://github.com/kangarooking/cangjie-skill/blob/main/cangjie.py) script.

### 1. Extract Transcripts with video-downloader

First, convert video content to text using the companion `video-downloader` skill referenced in the repository's [`README.md`](https://github.com/kangarooking/cangjie-skill/blob/main/README.md). This utility downloads the source video and either extracts embedded subtitle tracks (`.srt`, `.vtt`) or runs speech-to-text to generate a raw transcript file.

### 2. Clean and Strip Timestamps

Run a preprocessing script to remove timestamp lines and formatting artifacts. The pipeline requires plain text input to generate valid skill artifacts like [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) and [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md).

### 3. Execute the cangjie Pipeline

Feed the cleaned transcript to the main entry point:

```bash
python cangjie.py --input ./clean/video.txt --output ./skills

```

This command runs the Adler overview, parallel extractors, triple verification, RIA++ construction, Zettelkasten linking, and pressure-testing to emit final artifacts including [`DIGEST.md`](https://github.com/kangarooking/cangjie-skill/blob/main/DIGEST.md) and individual `*/SKILL.md` files.

## Python Utility for Timestamp Removal

Create a local helper script to sanitize subtitle files. This script removes lines matching the SRT/VTT timestamp pattern (`HH:MM:SS --> HH:MM:SS`) and strips index numbers and empty lines.

```python

# clean_subtitles.py – removes SRT/VTT timestamps and writes plain text

import re, sys, pathlib

TIMESTAMP_RE = re.compile(r'^\s*\d{2}:\d{2}:\d{2}(?:[.,]\d+)?\s*-->?\s*\d{2}:\d{2}:\d{2}(?:[.,]\d+)?\s*$')

def clean_subtitle_file(src_path: pathlib.Path, dst_path: pathlib.Path) -> None:
    """Read a subtitle file (SRT/VTT) and write a timestamp‑free plain‑text version."""
    with src_path.open(encoding="utf-8") as src, dst_path.open("w", encoding="utf-8") as dst:
        for line in src:
            if TIMESTAMP_RE.match(line):
                continue                     # skip timestamps

            if line.strip().isdigit():
                continue                     # skip index numbers (optional)

            if line.strip() == "":
                continue                     # skip empty lines

            dst.write(line.strip() + " ")

if __name__ == "__main__":
    if len(sys.argv) != 3:
        sys.exit("Usage: python clean_subtitles.py <input.srt> <output.txt>")
    clean_subtitle_file(pathlib.Path(sys.argv[1]), pathlib.Path(sys.argv[2]))

```

**Typical usage**:

```bash

# Download video & get subtitles

python video_downloader.py --url "https://www.youtube.com/watch?v=XYZ" --out-dir ./raw

# Clean the subtitle file

python clean_subtitles.py ./raw/video.srt ./clean/video.txt

# Run cangjie-skill on the cleaned transcript

python cangjie.py --input ./clean/video.txt --output ./skills

```

## Optional: Preserving Timestamp References

If traceability is required, store the original subtitle file alongside the cleaned version. You can map generated skill entries back to specific video timestamps by storing line range metadata in a JSON sidecar:

```json
{
  "skill_id": "example-skill",
  "source": {
    "subtitle_file": "video.srt",
    "line_range": "45-47"
  }
}

```

The generated [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) files can reference this mapping in their `source` fields, allowing users to locate the exact moment in the original video while keeping the pipeline's input free of temporal noise.

## Summary

- **Timestamps must be stripped** from video subtitles before processing with [`cangjie.py`](https://github.com/kangarooking/cangjie-skill/blob/main/cangjie.py) to avoid RIA-TV++ extractor failures.
- Use the companion `video-downloader` skill to extract `.srt` or `.vtt` files from video sources.
- Run a Python preprocessing script (like [`clean_subtitles.py`](https://github.com/kangarooking/cangjie-skill/blob/main/clean_subtitles.py)) to remove timestamp patterns, index numbers, and empty lines.
- Feed only the cleaned plain text to the cangjie pipeline to ensure valid [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md), [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md), and [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) generation.
- Retain original subtitle files for reference mapping if source traceability is required.

## Frequently Asked Questions

### What happens if I don't remove timestamps from video transcripts?

The RIA-TV++ extractors will treat timestamp strings as content tokens, causing the **triple verification** step to fail when it cannot find meaningful semantic citations. This results in empty or incomplete skill artifacts because the extractors cannot validate concepts against temporal metadata.

### Can cangjie-skill process audio podcasts the same way as videos?

Yes. Audio content requires the same preprocessing: transcribe the audio to text using a speech-to-text engine, strip any generated timestamps or speaker labels, and process the resulting plain text through [`cangjie.py`](https://github.com/kangarooking/cangjie-skill/blob/main/cangjie.py). The pipeline handles any long-form text regardless of original media type.

### Where does the cleaned text get stored in the cangjie output structure?

The cleaned input text is not stored in the output; rather, the pipeline generates structured artifacts in the `--output` directory, including [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) (high-level summary), [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md) (searchable master list), [`DIGEST.md`](https://github.com/kangarooking/cangjie-skill/blob/main/DIGEST.md) (condensed takeaways), and individual [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) files in subdirectories as defined in the repository's [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) schema specification.

### Is there an automated way to handle timestamp removal in CI/CD?

Yes. Integrate the [`clean_subtitles.py`](https://github.com/kangarooking/cangjie-skill/blob/main/clean_subtitles.py) script into your CI workflow before the cangjie execution step. The repository's `.github/workflows/` directory (containing examples like [`update-star-history.yml`](https://github.com/kangarooking/cangjie-skill/blob/main/update-star-history.yml)) demonstrates the project's automation patterns, which you can extend to include subtitle preprocessing as a prerequisite job before running the main skill generation pipeline.