# How to Contribute to the Development of cangjie-skill: A Complete Guide

> Learn how to contribute to cangjie-skill development. Add extractors, extend schemas, improve docs, create skill packs, or enhance CI. Explore our modular pipeline for parallel contributions.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: how-to-guide
- Published: 2026-07-17

---

**You can contribute to cangjie-skill by adding extractors in `extractors/`, extending the RIA++ schema in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md), improving methodology documentation, creating new skill packs, or enhancing CI automation—the pipeline is modular by design to support parallel contributions across all seven stages.**

The kangarooking/cangjie-skill repository transforms books, videos, and podcasts into AI-callable "skills" through a modular pipeline. If you want to contribute to the development of cangjie-skill, the architecture is deliberately designed to let you plug into any stage of the **RIA-TV++** workflow without disrupting other components. This guide covers the specific file paths, contribution patterns, and submission workflows used by the project.

## Understanding the Modular RIA-TV++ Architecture

The cangjie-skill pipeline processes raw content through seven distinct stages, from initial overview to final delivery. Each stage is defined in the `methodology/` directory and operates on specific artifacts.

### The Seven-Stage Pipeline

- **Stage 0 – Overview**: Defines project motivation and ecosystem in [`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md) and the root [`README.md`](https://github.com/kangarooking/cangjie-skill/blob/main/README.md).

- **Stage 1 – Adler Analysis**: Performs high-level reading of source text to produce [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md).

- **Stage 2 – Parallel Extraction**: Fires five specialized extractors defined in [`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md), [`extractors/principle-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/principle-extractor.md), [`extractors/case-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/case-extractor.md), [`extractors/counter-example-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/counter-example-extractor.md), and [`extractors/glossary-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/glossary-extractor.md).

- **Stage 3 – Triple Verification**: Validates candidates for citation strength, predictive power, and uniqueness as documented in [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md).

- **Stage 4 – RIA++ Construction**: Formats verified items into the six-field RIA++ schema defined in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md).

- **Stage 5 – Zettelkasten Linking**: Builds dependency indexes following [`methodology/05-stage3-zettelkasten.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/05-stage3-zettelkasten.md).

- **Stage 6 – Pressure Testing**: Generates bait-question test cases according to [`methodology/06-stage4-pressure-test.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/06-stage4-pressure-test.md).

- **Stage 7 – Delivery**: Emits [`DIGEST.md`](https://github.com/kangarooking/cangjie-skill/blob/main/DIGEST.md) and installs packs into Claude Code or Cursor per [`methodology/07-stage5-deliver.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/07-stage5-deliver.md).

All stages rely on **template files** in `templates/` and **prompt files** in `extractors/` to transform data into markdown artifacts.

## Six Ways to Contribute to cangjie-skill

### 1. Add or Improve Extractors

Create new prompt markdown files in `extractors/` or refine existing ones. The five core extractors (framework, principle, case, counter-example, glossary) run in parallel during Stage 2. New extractors must output JSON conforming to the schema in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md).

### 2. Extend the RIA++ Schema

Modify [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) to add new fields to the six-field RIA++ schema (R/I/A1/A2/E/B). Update corresponding templates in `templates/` to render the new fields into [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md), [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md), and [`DIGEST.md`](https://github.com/kangarooking/cangjie-skill/blob/main/DIGEST.md).

### 3. Enhance Methodology Documentation

Clarify existing stages by editing files in `methodology/`. For example, update the triple verification criteria in [`methodology/03-stage1.5-triple-verify.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/03-stage1.5-triple-verify.md) or add visual workflow diagrams to [`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md).

### 4. Create New Skill Packs

Generate complete skill directories containing [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md), [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md), and [`DIGEST.md`](https://github.com/kangarooking/cangjie-skill/blob/main/DIGEST.md), then submit them as new top-level directories. This requires no code changes—only properly formatted markdown artifacts.

### 5. Automate CI and Testing

The repository currently ships a GitHub Action for star-history generation. You can add unit-style checks for extractor prompt syntax validation or schema compliance tests in `scripts/`.

### 6. Documentation and Localization

Improve existing translations in [`README.en.md`](https://github.com/kangarooking/cangjie-skill/blob/main/README.en.md) or [`README.ja.md`](https://github.com/kangarooking/cangjie-skill/blob/main/README.ja.md), or add new language versions following the structure of the root [`README.md`](https://github.com/kangarooking/cangjie-skill/blob/main/README.md).

## Step-by-Step Contribution Workflows

### Running the Pipeline Locally

Test your changes end-to-end using the pipeline entry point. You need an OpenAI-compatible LLM endpoint configured in your environment.

```bash

# Install dependencies

pip install -r requirements.txt

# Prepare input directory

mkdir -p input && cp path/to/book.txt input/

# Execute pipeline

python scripts/run_pipeline.py input/book.txt output/

```

### Adding a New Extractor

Create a new extractor by adding a prompt file and registering it in the configuration.

```bash

# Create the prompt file

cat > extractors/strategy-extractor.md <<'EOF'

# Strategy Extractor Prompt

You are a knowledge engineer. Scan the provided text and output every distinct
strategy, its purpose, and a concrete example. Use the following JSON schema:
{
  "strategy": "<name>",
  "purpose": "<short description>",
  "example": "<quote or scenario>"
}
EOF

# Register in pipeline_config.yaml under the extractors: section

```

Ensure your extractor references the central [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) definition so downstream tools can consume the output.

### Submitting a New Skill Pack

Fork the repository, then use standard Git workflow to submit your generated skills.

```bash

# Create feature branch

git checkout -b new-skill-pack

# Copy generated artifacts

cp -r output/my-book my-book-skill/

# Commit and push

git add my-book-skill/
git commit -m "Add skill pack for My Book"
git push origin new-skill-pack

# Open Pull Request on kangarooking/cangjie-skill

```

## Key Files Every Contributor Must Know

These files constitute the "contract" that contributors extend:

- **[`README.md`](https://github.com/kangarooking/cangjie-skill/blob/main/README.md)** – Overall project introduction, contribution guide, and ecosystem overview.

- **[`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md)** – Formal definition of the RIA-TV++ skill schema used by downstream agents like Claude Code and Cursor.

- **`methodology/*.md`** – Detailed documentation for each pipeline stage, from [`00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/00-overview.md) to [`07-stage5-deliver.md`](https://github.com/kangarooking/cangjie-skill/blob/main/07-stage5-deliver.md).

- **`extractors/*.md`** – LLM prompt templates for the five core extraction types.

- **`templates/*.md.template`** – Jinja-style templates that render final artifacts.

Updating any of these files automatically propagates through the pipeline and improves all subsequently generated skill packs.

## Summary

- **cangjie-skill** uses a seven-stage RIA-TV++ pipeline defined in `methodology/` and driven by `extractors/` and `templates/`.
- **Extractors** are the primary extension point—add new ones by creating markdown files in `extractors/` and registering them in [`pipeline_config.yaml`](https://github.com/kangarooking/cangjie-skill/blob/main/pipeline_config.yaml).
- **Schema changes** require updating [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) and corresponding templates in `templates/`.
- **Skill packs** can be contributed as standalone directories without modifying core pipeline code.
- **All contributions** must reference the central [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) definition to ensure compatibility with downstream agents.

## Frequently Asked Questions

### What programming languages are required to contribute to cangjie-skill?

The pipeline primarily uses **Python** for automation scripts and **Markdown** for prompt engineering and documentation. Extractors are written as Markdown prompt files rather than code, making contributions accessible to domain experts without deep programming experience.

### Do I need an LLM API key to test my changes?

Yes. Running the full pipeline locally requires an **OpenAI-compatible LLM endpoint** configured in your environment. The [`scripts/run_pipeline.py`](https://github.com/kangarooking/cangjie-skill/blob/main/scripts/run_pipeline.py) entry point makes API calls to process raw transcripts through the extraction stages.

### How do I validate my new extractor before submitting?

Test your extractor against sample texts using the local pipeline command. Verify that output conforms to the JSON schema defined in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) and that the markdown renders correctly through the templates in `templates/`. Check that your prompt file syntax matches the style of existing files like [`extractors/framework-extractor.md`](https://github.com/kangarooking/cangjie-skill/blob/main/extractors/framework-extractor.md).

### Can I contribute without modifying the core pipeline code?

Yes. **Creating new skill packs** is a zero-code contribution path. Simply generate the required artifacts ([`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md), [`INDEX.md`](https://github.com/kangarooking/cangjie-skill/blob/main/INDEX.md), [`DIGEST.md`](https://github.com/kangarooking/cangjie-skill/blob/main/DIGEST.md)) using the existing pipeline, place them in a new directory, and submit via Pull Request. This approach is ideal for subject matter experts adding content rather than changing tooling.